0
Models tracked
0
Providers
0
With benchmarks
0
Open source
Model Rankings
Top models ranked by publicly reported benchmark score. Every number links to its source.
Top models per task
Best-in-class per task type, ranked by publicly reported benchmark scores. Rows with no verified score are omitted.
Coding
SWE-bench ProReasoning
GPQA DiamondAgents
OSWorldVision
MMMULong Context
Context windowBudget
Blended price per MTok, lower is betterBenchmark leaderboards
Top models per benchmark. Click any row to see full model specs.
Fastest and most affordable models
Median output speed as measured by third-party benchmarks (Artificial Analysis / OpenRouter); pricing as listed by the provider.
Model comparison
37 of 37 models. Sort by any column. Filter by provider, license, or task.
| Model ▼ | Provider ▼ | Open | Context ▼ | In $/MTok ▼ | Out $/MTok | Tok/s ▼ | Released ▼ | Best for |
|---|---|---|---|---|---|---|---|---|
| Anthropic | 1.0M | $4.00 | $20.00 | -- | 2026-09 | CodeAgentsReasoning | ||
| OpenAI | 1.1M | $2.00 | $10.00 | -- | 2026-09 | CodeAgentsReasoning | ||
| OpenAI | 1.1M | $0.10 | $0.50 | -- | 2026-09 | Budget | ||
| OpenAI | 1.1M | $10.00 | $50.00 | -- | 2026-09 | AgentsCoderesearch | ||
| 1.0M | $0.75 | $3.75 | -- | 2026-09 | CodeAgentsBudget | |||
| Anthropic | 1.0M | $10.00 | $50.00 | -- | 2026-09 | CodeAgentsReasoningresearch | ||
| Anthropic | 1.0M | $10.00 | $50.00 | -- | 2026-09 | CodeAgentsReasoningresearch | ||
| Anthropic | 1.0M | $5.00 | $25.00 | 52.3 | 2026-07 | CodeAgentsReasoning | ||
| 2.1M | -- | -- | -- | 2026-07 | Long CtxReasoningVision | |||
| OpenAI | 1.1M | $4.00 | $20.00 | -- | 2026-07 | CodeAgentsReasoning | ||
| Tencent | ✓ | 262K | $0.20 | $0.80 | -- | 2026-07 | CodeReasoningAgents | |
| Anthropic | 1.0M | $2.00 | $10.00 | 79 | 2026-06 | CodeAgentsReasoning | ||
| Sakana AI | 1.0M | $5.00 | $30.00 | -- | 2026-06 | ReasoningAgentsCode | ||
| Z.ai | ✓ | 1.0M | $0.90 | $2.86 | -- | 2026-06 | CodeReasoningAgents | |
| Anthropic | 1.0M | $10.00 | $50.00 | 70 | 2026-06 | CodeAgentsReasoning | ||
| Anthropic | 1.0M | $5.00 | $25.00 | 65 | 2026-05 | CodeAgentsReasoning | ||
| 1.0M | $1.50 | $9.00 | 198 | 2026-05 | CodeAgentsBudget | |||
| DeepSeek | ✓ | 1.0M | $0.43 | $0.87 | 72 | 2026-04 | CodeReasoningBudget | |
| DeepSeek | ✓ | 1.0M | $0.09 | $0.18 | 113 | 2026-04 | BudgetCode | |
| OpenAI | 1.0M | $5.00 | $30.00 | 89 | 2026-04 | CodeReasoningAgents | ||
| Alibaba | ✓ | 1.0M | $1.25 | $3.75 | 206 | 2026-04 | ReasoningCodeLong Ctx | |
| xAI | 2.0M | $2.00 | $6.00 | -- | 2026-03 | ReasoningAgentsLong Ctx | ||
| OpenAI | 1.0M | $2.50 | $15.00 | 165 | 2026-03 | CodeReasoningVision | ||
| Anthropic | 1.0M | $3.00 | $15.00 | 47 | 2026-02 | CodeAgentsVision | ||
| Anthropic | 200K | $1.00 | $5.00 | 98 | 2025-10 | BudgetAgents | ||
| Mistral | 128K | $0.10 | $0.30 | 166 | 2025-09 | BudgetCode | ||
| OpenAI | 400K | $0.63 | $5.00 | 107 | 2025-08 | ReasoningCode | ||
| 1.0M | $1.25 | $10.00 | 151 | 2025-06 | Long CtxReasoningVision | |||
| 1.0M | $0.30 | $2.50 | 222 | 2025-05 | BudgetCodeAgents | |||
| Alibaba | ✓ | 262K | $0.78 | $3.90 | 66 | 2025-04 | ReasoningCode | |
| OpenAI | 200K | $1.10 | $4.40 | 163 | 2025-04 | ReasoningCodeBudget | ||
| OpenAI | 200K | $2.00 | $8.00 | 167 | 2025-04 | ReasoningCode | ||
| Meta | ✓ | 1.0M | $0.15 | $0.60 | 118 | 2025-04 | VisionReasoningBudget | |
| Meta | ✓ | 10.0M | $0.08 | $0.30 | 109 | 2025-04 | Long CtxBudgetVision | |
| xAI | 131K | $3.00 | $15.00 | -- | 2025-02 | ReasoningCode | ||
| DeepSeek | ✓ | 128K | $0.55 | $2.19 | -- | 2025-01 | ReasoningCode | |
| Mistral | 128K | $2.00 | $6.00 | 55 | 2024-07 | CodeReasoning |
Context window, cost and speed
Full specs side by side. Click a column header to sort.
| Model ▼ | Provider | Context ▼ | Max output | In $/MTok ▼ | Out $/MTok | Tok/s ▼ | Pricing note |
|---|---|---|---|---|---|---|---|
| Meta | 10.0M | -- | $0.08 | $0.30 | 109 | Via third-party inference providers; Meta does not operate a paid API | |
| 2.1M | -- | -- | -- | -- | GA expected Jul 17 2026; pricing not confirmed as of 2026-07-06 | ||
| xAI | 2.0M | -- | $2.00 | $6.00 | -- | ||
| OpenAI | 1.1M | 128K | $10.00 | $50.00 | -- | Cached input $1/MTok, cache writes $12.50/MTok. Prompts over 272K input tokens are billed at 2x the input/cache rates and 1.5x the output rate for the entire request, not just the excess. | |
| OpenAI | 1.1M | 128K | $2.00 | $10.00 | -- | Cached input $0.20/MTok. Prompts over 272K input tokens cost 2x input and cache rates and 1.5x output for the full request. Source: developers.openai.com/api/docs/models/gpt-6-sol, checked 2026-09-26. | |
| OpenAI | 1.1M | 128K | $0.10 | $0.50 | -- | Cached input $0.01/MTok. Prompts over 272K input tokens cost 2x input and cache rates and 1.5x output. Source: developers.openai.com/api/docs/models/gpt-6-luna, checked 2026-09-26. | |
| OpenAI | 1.1M | 128K | $4.00 | $20.00 | -- | OpenAI cut pricing from launch $5/$30 to $4/$20 per MTok on Aug 21 2026, a promotion running at least through Nov 21 2026. Prompts over 272K input tokens are billed at 2x the input/cache rates and 1.5x the output rate for the entire request, not just the excess; cache writes are 1.25x the uncached input rate. | |
| Anthropic | 1.0M | 128K | $10.00 | $50.00 | -- | Cache reads cut to $0.25/MTok, a 75% reduction from Fable 5's $1.00/MTok. 5-minute cache writes $12.50/MTok, 1-hour cache writes $20/MTok. Batch API is 50% off. Anthropic reports this brings typical-workload cost down about 25% and highly agentic workload cost down up to about 45% versus Fable 5 at the same prices. | |
| Anthropic | 1.0M | 128K | $10.00 | $50.00 | -- | Identical pricing to Claude Fable 5.1: cache reads $0.25/MTok (2.5% of base input), 5-minute cache writes $12.50/MTok, 1-hour cache writes $20/MTok, Batch API 50% off. | |
| Anthropic | 1.0M | 128K | $4.00 | $20.00 | -- | Cache reads $0.20/MTok (5% of input), 5-min cache writes $5, 1-hour $8, Batch API 50% off. Source: platform.claude.com/docs/en/about-claude/pricing, checked 2026-09-26. | |
| Anthropic | 1.0M | -- | $5.00 | $25.00 | 52.3 | ||
| Anthropic | 1.0M | 128K | $10.00 | $50.00 | 70 | ||
| Anthropic | 1.0M | 128K | $2.00 | $10.00 | 79 | Launch intro price of $2/$10 is now the standard price; the planned Sep 1 2026 rise to $3/$15 will not occur. Source: platform.claude.com/docs/en/about-claude/pricing, checked 2026-09-26. | |
| Anthropic | 1.0M | 128K | $5.00 | $25.00 | 65 | ||
| Anthropic | 1.0M | 128K | $3.00 | $15.00 | 47 | ||
| OpenAI | 1.0M | -- | $5.00 | $30.00 | 89 | ||
| OpenAI | 1.0M | 128K | $2.50 | $15.00 | 165 | ||
| 1.0M | 66K | $1.50 | $9.00 | 198 | |||
| 1.0M | 66K | $0.75 | $3.75 | -- | $0.75/$3.75 through Dec 31 2026, then $1.50/$7.50 from Jan 1 2027 (same schedule as Gemini 3.7 Flash). Source: ai.google.dev/gemini-api/docs/pricing, checked 2026-09-26. | ||
| 1.0M | -- | $1.25 | $10.00 | 151 | $2.50/$15 per MTok for prompts over 200K tokens | ||
| 1.0M | 66K | $0.30 | $2.50 | 222 | |||
| Z.ai | 1.0M | 131K | $0.90 | $2.86 | -- | Via OpenRouter (model ID z-ai/glm-5.2); SiliconFlow lists $1.40/$4.40 per MTok | |
| Meta | 1.0M | -- | $0.15 | $0.60 | 118 | Via third-party inference providers; Meta does not operate a paid API | |
| DeepSeek | 1.0M | 384K | $0.43 | $0.87 | 72 | 75% launch promo made permanent 2026-05-22 (original list was $1.74/$3.48 per MTok); $0.435 is now the standard rate. | |
| DeepSeek | 1.0M | 384K | $0.09 | $0.18 | 113 | Cached input $0.0028/MTok | |
| Alibaba | 1.0M | -- | $1.25 | $3.75 | 206 | ||
| Sakana AI | 1.0M | 131K | $5.00 | $30.00 | -- | $10/$45 per MTok for contexts over 272K tokens; cached input $0.50/MTok (standard), $1.00/MTok (over 272K) | |
| OpenAI | 400K | -- | $0.63 | $5.00 | 107 | Succeeded by GPT-5.5 (official OpenAI replacement); no longer on pricing page. API shutdown Dec 11 2026. | |
| Alibaba | 262K | -- | $0.78 | $3.90 | 66 | ||
| Tencent | 262K | -- | $0.20 | $0.80 | -- | Free via OpenRouter for ~2 weeks from 2026-07-07 (model ID tencent/hy3:free); paid tier $0.20/$0.80 per MTok via OpenRouter | |
| Anthropic | 200K | -- | $1.00 | $5.00 | 98 | ||
| OpenAI | 200K | -- | $1.10 | $4.40 | 163 | ||
| OpenAI | 200K | -- | $2.00 | $8.00 | 167 | ||
| xAI | 131K | -- | $3.00 | $15.00 | -- | ||
| DeepSeek | 128K | -- | $0.55 | $2.19 | -- | deepseek-reasoner API alias now routes to V4 Flash thinking mode | |
| Mistral | 128K | -- | $2.00 | $6.00 | 55 | ||
| Mistral | 128K | -- | $0.10 | $0.30 | 166 |
Benchmark glossary
What each benchmark actually measures and why it matters.
SWE-bench Verified
Real GitHub issues from popular Python repos. The model must write code that passes the existing test suite. "Verified" means a human checked that each issue is solvable. Scores range 0-100%.
swebench.comMMLU
Massive Multitask Language Understanding. 57 academic subjects from high-school to professional level (law, medicine, STEM). Tests breadth of knowledge. Scores are % correct.
Hendrycks et al.HumanEval
OpenAI benchmark: 164 hand-crafted Python programming problems. Measures functional code generation (pass@1). Widely used but considered dated for frontier models.
OpenAIMMMU
Massive Multidisciplinary Multimodal Understanding. 11.5K questions requiring images (charts, diagrams, photos) plus text. Standard for vision-language model evaluation.
mmmu-benchmark.github.ioRULER
Ruler for Long-Context Evaluation. Tests retrieval, multi-hop reasoning, and aggregation over very long documents (up to 128K tokens). Better signal than "needle in haystack" for long-context claims.
Hsieh et al.GPQA Diamond
Graduate-Level Google-Proof Q&A, Diamond split. Expert-level science questions that Google cannot directly answer. High signal for deep scientific reasoning. Human experts score ~65%.
Rein et al.MATH
Competition mathematics (AMC, AIME, Olympiad). 12,500 problems across 7 difficulty levels. Tests symbolic reasoning and multi-step derivation. Scores are % correct.
Hendrycks et al.Not sure which model fits your project?
We evaluate models against your actual use case, run cost-to-quality tests across providers, and build the integration. No vendor lock-in, no guesswork.
Frequently asked questions
What is an LLM leaderboard?
An LLM leaderboard ranks AI language models by benchmark scores, speed, pricing, and context window so developers can pick the best model for their use case without reading every provider's documentation separately.
Which AI model scores highest on benchmarks in 2026?
Claude Fable 5 leads on SWE-bench Verified (95.0%) for coding. GPT-5-5 and Gemini 2.5 Pro are competitive on general reasoning. Use the filter above to sort by the benchmark most relevant to your workload.
How do I compare AI models side by side?
Use the leaderboard filters on this page, or go to the Compare tool to select up to four models and view their specs and benchmarks in a single table.
Which AI model API is the cheapest?
Open-weight models like Llama 4 Scout and Qwen3 are available free through several providers. Among paid APIs, DeepSeek V4 Pro offers strong reasoning at a lower cost per token than GPT-5 or Claude Fable 5. See the full pricing breakdown at /models/pricing.
Related guides
How we verify. Benchmark scores are taken from official provider pages, third-party leaderboards, or peer-reviewed papers. Where a number could not be independently confirmed it is shown as a dash. Pricing reflects the public API rate at the time of last update; check the provider for current pricing.
Affiliate disclosure. AY Automate has no affiliate relationship with any model provider listed here. Rankings are editorial, not commercial.