Skip to content
AI models under $10 / 1M output tokens

Set the budget first. Then find the strongest model that survives it.

A hard output-price ceiling is more useful than a vague request for a “cheap model.” ModelShortlist can filter current OpenRouter pricing first, then let the host AI compare capability and benchmark evidence among eligible models.

The decision

The answer changes with the workload.

Budget-constrained model selection should fail closed on the constraint rather than recommending an attractive model that only looks cheap at one pricing tier or provider. Once the ceiling is enforced, the remaining decision depends on workload quality, tools, context, and how much model performance you are willing to trade for lower cost.

Current output-token pricing across applicable pricing tiers rather than a stale published estimate.

Minimum context and tool requirements so the cheapest model is still operationally usable.

Artificial Analysis quality evidence where the model identity matches confidently.

Whether the workload is high-volume enough that small token-price differences dominate the economics.

Independent performance evidence

Artificial Analysis

ModelShortlist uses Artificial Analysis benchmark and performance evidence when the model identity can be reconciled confidently. It does not create, relabel, or pretend ownership of those benchmarks.

How Artificial Analysis contributes

Current operational facts

OpenRouter

The current catalog supplies model availability, context, supported parameters, pricing, and provider details. ZDR endpoint evidence is applied only when the workload explicitly requires ZDR.

Compare OpenRouter models by workload

Model selection is time-sensitive: new models launch, prices move, benchmark results change, tool support evolves, and providers add or remove endpoints. ModelShortlist surfaces freshness and degraded upstream state instead of silently presenting stale evidence as current.

Ask naturally

Prompts that work.

ModelShortlist sits behind the AI assistant you already use. Describe the job and hard constraints instead of translating them into a fixed ranking formula.

I need the strongest current coding model under $10 per million output tokens, with tool calling and at least 100k context.
What is the best-value model for this workload if $5 per million output tokens is a hard ceiling?
Compare the eligible models under my output-price ceiling and tell me what quality I give up versus the unrestricted shortlist.

Why ModelShortlist

Current evidence, not a static leaderboard.

Uses current OpenRouter pricing instead of a static cost table.

Applies price ceilings conservatively across tiered pricing rather than accepting a model whose higher tier exceeds the limit.

Keeps capability and context constraints hard while letting the host AI reason about softer quality-versus-cost tradeoffs.

Adds Artificial Analysis evidence when available so “cheap” does not become a synonym for “good enough” without evidence.

Local BYOK architecture. No ModelShortlist account, hosted key vault, or telemetry in the MCP.

Put current model-selection evidence inside your assistant.

Install the local MCP, add your OpenRouter and Artificial Analysis keys, and ask the model-selection question in normal language.

Open the install configurator