Skip to content
AI models for coding agents

Choose a coding model for the agent you are actually running.

Autonomous coding changes the model-selection problem. Tool reliability, context, output cost, agentic performance, and long-running economics can matter more than a single benchmark score.

The decision

The answer changes with the workload.

A model that is excellent for a short interactive coding question may be a poor choice for dozens of parallel subagents or a long-running repository task. The right shortlist depends on how much autonomy you need, how often tools are called, how large the working context becomes, and how much output the agent generates.

Tool/function-calling support and the reliability needed for repeated agent actions.

Context capacity for repository state, plans, tool results, and long-running histories.

Coding and agentic benchmark evidence when a verified Artificial Analysis match is available.

Input/output pricing, especially when many subagents or long traces multiply token usage.

Independent performance evidence

Artificial Analysis

ModelShortlist uses Artificial Analysis benchmark and performance evidence when the model identity can be reconciled confidently. It does not create, relabel, or pretend ownership of those benchmarks.

How Artificial Analysis contributes

Current operational facts

OpenRouter

The current catalog supplies model availability, context, supported parameters, pricing, and provider details. ZDR endpoint evidence is applied only when the workload explicitly requires ZDR.

Compare OpenRouter models by workload

Model selection is time-sensitive: new models launch, prices move, benchmark results change, tool support evolves, and providers add or remove endpoints. ModelShortlist surfaces freshness and degraded upstream state instead of silently presenting stale evidence as current.

Ask naturally

Prompts that work.

ModelShortlist sits behind the AI assistant you already use. Describe the job and hard constraints instead of translating them into a fixed ranking formula.

I need the best-value model for a long-running autonomous coding agent. Tool calling is required and I need at least 100k context.
What is the cheapest OpenRouter model I would trust with repetitive coding subagents while preserving reliable tool use?
Quality matters more than cost for this repository refactor, but I still want to know whether the frontier premium is justified.

Why ModelShortlist

Current evidence, not a static leaderboard.

Starts from the current OpenRouter model catalog rather than a fixed list of fashionable coding models.

Can enforce tool calling, minimum context, creator filters, and price ceilings as hard constraints.

Attaches Artificial Analysis coding, intelligence, and agentic metrics only when the model identity can be matched confidently.

Leaves the final weighting to the host AI so an autonomous agent, cheap subagent, and interactive coding assistant can produce different recommendations.

Local BYOK architecture. No ModelShortlist account, hosted key vault, or telemetry in the MCP.

Put current model-selection evidence inside your assistant.

Install the local MCP, add your OpenRouter and Artificial Analysis keys, and ask the model-selection question in normal language.

Open the install configurator