0
Models tracked
0
Providers
0
With benchmarks
0
Open source
Model Rankings
Top models ranked by publicly reported benchmark score. Every number links to its source.
Top models per task
Best-in-class per task type, ranked by publicly reported benchmark scores. Rows with no verified score are omitted.
Coding
SWE-bench ProReasoning
GPQA DiamondAgents
OSWorldVision
MMMULong Context
Context windowBudget
Blended price per MTok, lower is betterBenchmark leaderboards
Top models per benchmark. Click any row to see full model specs.
Fastest and most affordable models
Median output speed as measured by third-party benchmarks (Artificial Analysis / OpenRouter); pricing as listed by the provider.
Model comparison
28 of 28 models. Sort by any column. Filter by provider, license, or task.
| Model ▼ | Provider ▼ | Open | Context ▼ | In $/MTok ▼ | Out $/MTok | Tok/s ▼ | Released ▼ | Best for |
|---|---|---|---|---|---|---|---|---|
| 2.1M | -- | -- | -- | 2026-07 | Long CtxReasoningVision | |||
| Tencent | ✓ | 262K | $0.20 | $0.80 | -- | 2026-07 | CodeReasoningAgents | |
| Anthropic | 1.0M | $2.00 | $10.00 | 79 | 2026-06 | CodeAgentsReasoning | ||
| Sakana AI | 1.0M | $5.00 | $30.00 | -- | 2026-06 | ReasoningAgentsCode | ||
| Z.ai | ✓ | 1.0M | $0.90 | $2.86 | -- | 2026-06 | CodeReasoningAgents | |
| Anthropic | 1.0M | $10.00 | $50.00 | 70 | 2026-06 | CodeAgentsReasoning | ||
| Anthropic | 1.0M | $5.00 | $25.00 | 65 | 2026-05 | CodeAgentsReasoning | ||
| 1.0M | $1.50 | $9.00 | 198 | 2026-05 | CodeAgentsBudget | |||
| DeepSeek | ✓ | 1.0M | $0.43 | $0.87 | 72 | 2026-04 | CodeReasoningBudget | |
| DeepSeek | ✓ | 1.0M | $0.09 | $0.18 | 113 | 2026-04 | BudgetCode | |
| OpenAI | 1.0M | $5.00 | $30.00 | 89 | 2026-04 | CodeReasoningAgents | ||
| Alibaba | ✓ | 1.0M | $1.25 | $3.75 | 206 | 2026-04 | ReasoningCodeLong Ctx | |
| xAI | 2.0M | $2.00 | $6.00 | -- | 2026-03 | ReasoningAgentsLong Ctx | ||
| OpenAI | 1.0M | $2.50 | $15.00 | 165 | 2026-03 | CodeReasoningVision | ||
| Anthropic | 1.0M | $3.00 | $15.00 | 47 | 2026-02 | CodeAgentsVision | ||
| Anthropic | 200K | $1.00 | $5.00 | 98 | 2025-10 | BudgetAgents | ||
| Mistral | 128K | $0.10 | $0.30 | 166 | 2025-09 | BudgetCode | ||
| OpenAI | 400K | $0.63 | $5.00 | 107 | 2025-08 | ReasoningCode | ||
| 1.0M | $1.25 | $10.00 | 151 | 2025-06 | Long CtxReasoningVision | |||
| 1.0M | $0.30 | $2.50 | 222 | 2025-05 | BudgetCodeAgents | |||
| Alibaba | ✓ | 262K | $0.78 | $3.90 | 66 | 2025-04 | ReasoningCode | |
| OpenAI | 200K | $1.10 | $4.40 | 163 | 2025-04 | ReasoningCodeBudget | ||
| OpenAI | 200K | $2.00 | $8.00 | 167 | 2025-04 | ReasoningCode | ||
| Meta | ✓ | 1.0M | $0.15 | $0.60 | 118 | 2025-04 | VisionReasoningBudget | |
| Meta | ✓ | 10.0M | $0.08 | $0.30 | 109 | 2025-04 | Long CtxBudgetVision | |
| xAI | 131K | $3.00 | $15.00 | -- | 2025-02 | ReasoningCode | ||
| DeepSeek | ✓ | 128K | $0.55 | $2.19 | -- | 2025-01 | ReasoningCode | |
| Mistral | 128K | $2.00 | $6.00 | 55 | 2024-07 | CodeReasoning |
Context window, cost and speed
Full specs side by side. Click a column header to sort.
| Model ▼ | Provider | Context ▼ | Max output | In $/MTok ▼ | Out $/MTok | Tok/s ▼ | Pricing note |
|---|---|---|---|---|---|---|---|
| Meta | 10.0M | -- | $0.08 | $0.30 | 109 | Via third-party inference providers; Meta does not operate a paid API | |
| 2.1M | -- | -- | -- | -- | GA expected Jul 17 2026; pricing not confirmed as of 2026-07-06 | ||
| xAI | 2.0M | -- | $2.00 | $6.00 | -- | ||
| Anthropic | 1.0M | 128K | $10.00 | $50.00 | 70 | ||
| Anthropic | 1.0M | 128K | $2.00 | $10.00 | 79 | Introductory pricing through Aug 31 2026; standard rate $3/$15 per MTok | |
| Anthropic | 1.0M | 128K | $5.00 | $25.00 | 65 | ||
| Anthropic | 1.0M | 128K | $3.00 | $15.00 | 47 | ||
| OpenAI | 1.0M | -- | $5.00 | $30.00 | 89 | ||
| OpenAI | 1.0M | 128K | $2.50 | $15.00 | 165 | ||
| 1.0M | 66K | $1.50 | $9.00 | 198 | |||
| 1.0M | -- | $1.25 | $10.00 | 151 | $2.50/$15 per MTok for prompts over 200K tokens | ||
| 1.0M | 66K | $0.30 | $2.50 | 222 | |||
| Z.ai | 1.0M | 131K | $0.90 | $2.86 | -- | Via OpenRouter (model ID z-ai/glm-5.2); SiliconFlow lists $1.40/$4.40 per MTok | |
| Meta | 1.0M | -- | $0.15 | $0.60 | 118 | Via third-party inference providers; Meta does not operate a paid API | |
| DeepSeek | 1.0M | 384K | $0.43 | $0.87 | 72 | 75% launch promo made permanent 2026-05-22 (original list was $1.74/$3.48 per MTok); $0.435 is now the standard rate. | |
| DeepSeek | 1.0M | 384K | $0.09 | $0.18 | 113 | Cached input $0.0028/MTok | |
| Alibaba | 1.0M | -- | $1.25 | $3.75 | 206 | ||
| Sakana AI | 1.0M | 131K | $5.00 | $30.00 | -- | $10/$45 per MTok for contexts over 272K tokens; cached input $0.50/MTok (standard), $1.00/MTok (over 272K) | |
| OpenAI | 400K | -- | $0.63 | $5.00 | 107 | Succeeded by GPT-5.5 (official OpenAI replacement); no longer on pricing page. API shutdown Dec 11 2026. | |
| Alibaba | 262K | -- | $0.78 | $3.90 | 66 | ||
| Tencent | 262K | -- | $0.20 | $0.80 | -- | Free via OpenRouter for ~2 weeks from 2026-07-07 (model ID tencent/hy3:free); paid tier $0.20/$0.80 per MTok via OpenRouter | |
| Anthropic | 200K | -- | $1.00 | $5.00 | 98 | ||
| OpenAI | 200K | -- | $1.10 | $4.40 | 163 | ||
| OpenAI | 200K | -- | $2.00 | $8.00 | 167 | ||
| xAI | 131K | -- | $3.00 | $15.00 | -- | ||
| DeepSeek | 128K | -- | $0.55 | $2.19 | -- | deepseek-reasoner API alias now routes to V4 Flash thinking mode | |
| Mistral | 128K | -- | $2.00 | $6.00 | 55 | ||
| Mistral | 128K | -- | $0.10 | $0.30 | 166 |
Benchmark glossary
What each benchmark actually measures and why it matters.
SWE-bench Verified
Real GitHub issues from popular Python repos. The model must write code that passes the existing test suite. "Verified" means a human checked that each issue is solvable. Scores range 0-100%.
swebench.comMMLU
Massive Multitask Language Understanding. 57 academic subjects from high-school to professional level (law, medicine, STEM). Tests breadth of knowledge. Scores are % correct.
Hendrycks et al.HumanEval
OpenAI benchmark: 164 hand-crafted Python programming problems. Measures functional code generation (pass@1). Widely used but considered dated for frontier models.
OpenAIMMMU
Massive Multidisciplinary Multimodal Understanding. 11.5K questions requiring images (charts, diagrams, photos) plus text. Standard for vision-language model evaluation.
mmmu-benchmark.github.ioRULER
Ruler for Long-Context Evaluation. Tests retrieval, multi-hop reasoning, and aggregation over very long documents (up to 128K tokens). Better signal than "needle in haystack" for long-context claims.
Hsieh et al.GPQA Diamond
Graduate-Level Google-Proof Q&A, Diamond split. Expert-level science questions that Google cannot directly answer. High signal for deep scientific reasoning. Human experts score ~65%.
Rein et al.MATH
Competition mathematics (AMC, AIME, Olympiad). 12,500 problems across 7 difficulty levels. Tests symbolic reasoning and multi-step derivation. Scores are % correct.
Hendrycks et al.Not sure which model fits your project?
We evaluate models against your actual use case, run cost-to-quality tests across providers, and build the integration. No vendor lock-in, no guesswork.
Frequently asked questions
What is an LLM leaderboard?
An LLM leaderboard ranks AI language models by benchmark scores, speed, pricing, and context window so developers can pick the best model for their use case without reading every provider's documentation separately.
Which AI model scores highest on benchmarks in 2026?
Claude Fable 5 leads on SWE-bench Verified (95.0%) for coding. GPT-5-5 and Gemini 2.5 Pro are competitive on general reasoning. Use the filter above to sort by the benchmark most relevant to your workload.
How do I compare AI models side by side?
Use the leaderboard filters on this page, or go to the Compare tool to select up to four models and view their specs and benchmarks in a single table.
Which AI model API is the cheapest?
Open-weight models like Llama 4 Scout and Qwen3 are available free through several providers. Among paid APIs, DeepSeek V4 Pro offers strong reasoning at a lower cost per token than GPT-5 or Claude Fable 5. See the full pricing breakdown at /models/pricing.
Related guides
How we verify. Benchmark scores are taken from official provider pages, third-party leaderboards, or peer-reviewed papers. Where a number could not be independently confirmed it is shown as a dash. Pricing reflects the public API rate at the time of last update; check the provider for current pricing.
Affiliate disclosure. AY Automate has no affiliate relationship with any model provider listed here. Rankings are editorial, not commercial.