Leaderboard

LLM Leaderboard 2026

AI model comparison across 37 models: benchmark scores, context windows, pricing, and speed. Every number comes from a public source or provider announcement -- no guesses, no missing data marked as zero. Filter by provider, task, or license to find the best LLM for your workload. Updated as new models ship.

0

Models tracked

0

Providers

0

With benchmarks

0

Open source

Model Rankings

Top models ranked by publicly reported benchmark score. Every number links to its source.

Top models per task

Best-in-class per task type, ranked by publicly reported benchmark scores. Rows with no verified score are omitted.

Coding

SWE-bench Pro

Reasoning

GPQA Diamond

Agents

OSWorld

Vision

MMMU

Long Context

Context window

Budget

Blended price per MTok, lower is better

Benchmark leaderboards

Top models per benchmark. Click any row to see full model specs.

ExploitBench (cybersecurity)

1GPT-6 Astra100%
2GPT-5.6 Sol78.5%

ScreenSpot-Pro (desktop automation)

1GPT-6 Astra92.7%
2GPT-5.6 Sol76.9%

Fastest and most affordable models

Median output speed as measured by third-party benchmarks (Artificial Analysis / OpenRouter); pricing as listed by the provider.

Fastest (tokens per second)

ModelProviderTok/sIn $/MTok
Gemini 2.5 FlashGoogle222$0.30
Qwen3.7 MaxAlibaba206$1.25
Gemini 3.5 FlashGoogle198$1.50
o3OpenAI167$2.00
Mistral Small 3Mistral166$0.10
GPT-5.4OpenAI165$2.50
o4-miniOpenAI163$1.10

Most affordable (input cost)

ModelProviderIn $/MTokOut $/MTok
Llama 4 ScoutMeta$0.08$0.30
DeepSeek V4 FlashDeepSeek$0.09$0.18
GPT-6 LunaOpenAI$0.10$0.50
Mistral Small 3Mistral$0.10$0.30
Llama 4 MaverickMeta$0.15$0.60
Tencent Hy3Tencent$0.20$0.80
Gemini 2.5 FlashGoogle$0.30$2.50

Model comparison

37 of 37 models. Sort by any column. Filter by provider, license, or task.

Model ▼Provider ▼OpenContext ▼In $/MTok ▼Out $/MTokTok/s ▼Released ▼Best for
Claude Opus 5.5Anthropic1.0M$4.00$20.00--2026-09
CodeAgentsReasoning
GPT-6 SolOpenAI1.1M$2.00$10.00--2026-09
CodeAgentsReasoning
GPT-6 LunaOpenAI1.1M$0.10$0.50--2026-09
Budget
GPT-6 AstraOpenAI1.1M$10.00$50.00--2026-09
AgentsCoderesearch
Gemini 3.8 FlashGoogle1.0M$0.75$3.75--2026-09
CodeAgentsBudget
Claude Fable 5.1Anthropic1.0M$10.00$50.00--2026-09
CodeAgentsReasoningresearch
Claude Mythos 5.1Anthropic1.0M$10.00$50.00--2026-09
CodeAgentsReasoningresearch
Claude Opus 5Anthropic1.0M$5.00$25.0052.32026-07
CodeAgentsReasoning
Gemini 3.5 ProGoogle2.1M------2026-07
Long CtxReasoningVision
GPT-5.6 SolOpenAI1.1M$4.00$20.00--2026-07
CodeAgentsReasoning
Tencent Hy3FreeTencent✓262K$0.20$0.80--2026-07
CodeReasoningAgents
Claude Sonnet 5Anthropic1.0M$2.00$10.00792026-06
CodeAgentsReasoning
Sakana Fugu UltraSakana AI1.0M$5.00$30.00--2026-06
ReasoningAgentsCode
GLM-5.2Z.ai✓1.0M$0.90$2.86--2026-06
CodeReasoningAgents
Claude Fable 5Anthropic1.0M$10.00$50.00702026-06
CodeAgentsReasoning
Claude Opus 4.8Anthropic1.0M$5.00$25.00652026-05
CodeAgentsReasoning
Gemini 3.5 FlashGoogle1.0M$1.50$9.001982026-05
CodeAgentsBudget
DeepSeek V4 ProDeepSeek✓1.0M$0.43$0.87722026-04
CodeReasoningBudget
DeepSeek V4 FlashDeepSeek✓1.0M$0.09$0.181132026-04
BudgetCode
GPT-5.5OpenAI1.0M$5.00$30.00892026-04
CodeReasoningAgents
Qwen3.7 MaxAlibaba✓1.0M$1.25$3.752062026-04
ReasoningCodeLong Ctx
Grok 4.20xAI2.0M$2.00$6.00--2026-03
ReasoningAgentsLong Ctx
GPT-5.4OpenAI1.0M$2.50$15.001652026-03
CodeReasoningVision
Claude Sonnet 4.6Anthropic1.0M$3.00$15.00472026-02
CodeAgentsVision
Claude Haiku 4.5Anthropic200K$1.00$5.00982025-10
BudgetAgents
Mistral Small 3Mistral128K$0.10$0.301662025-09
BudgetCode
GPT-5OpenAI400K$0.63$5.001072025-08
ReasoningCode
Gemini 2.5 ProGoogle1.0M$1.25$10.001512025-06
Long CtxReasoningVision
Gemini 2.5 FlashGoogle1.0M$0.30$2.502222025-05
BudgetCodeAgents
Qwen3 MaxAlibaba✓262K$0.78$3.90662025-04
ReasoningCode
o4-miniOpenAI200K$1.10$4.401632025-04
ReasoningCodeBudget
o3OpenAI200K$2.00$8.001672025-04
ReasoningCode
Llama 4 MaverickMeta✓1.0M$0.15$0.601182025-04
VisionReasoningBudget
Llama 4 ScoutMeta✓10.0M$0.08$0.301092025-04
Long CtxBudgetVision
Grok 3xAI131K$3.00$15.00--2025-02
ReasoningCode
DeepSeek R1DeepSeek✓128K$0.55$2.19--2025-01
ReasoningCode
Mistral Large 2Mistral128K$2.00$6.00552024-07
CodeReasoning

Context window, cost and speed

Full specs side by side. Click a column header to sort.

Model ▼ProviderContext ▼Max outputIn $/MTok ▼Out $/MTokTok/s ▼Pricing note
Llama 4 ScoutMeta10.0M--$0.08$0.30109Via third-party inference providers; Meta does not operate a paid API
Gemini 3.5 ProGoogle2.1M--------GA expected Jul 17 2026; pricing not confirmed as of 2026-07-06
Grok 4.20xAI2.0M--$2.00$6.00--
GPT-6 AstraOpenAI1.1M128K$10.00$50.00--Cached input $1/MTok, cache writes $12.50/MTok. Prompts over 272K input tokens are billed at 2x the input/cache rates and 1.5x the output rate for the entire request, not just the excess.
GPT-6 SolOpenAI1.1M128K$2.00$10.00--Cached input $0.20/MTok. Prompts over 272K input tokens cost 2x input and cache rates and 1.5x output for the full request. Source: developers.openai.com/api/docs/models/gpt-6-sol, checked 2026-09-26.
GPT-6 LunaOpenAI1.1M128K$0.10$0.50--Cached input $0.01/MTok. Prompts over 272K input tokens cost 2x input and cache rates and 1.5x output. Source: developers.openai.com/api/docs/models/gpt-6-luna, checked 2026-09-26.
GPT-5.6 SolOpenAI1.1M128K$4.00$20.00--OpenAI cut pricing from launch $5/$30 to $4/$20 per MTok on Aug 21 2026, a promotion running at least through Nov 21 2026. Prompts over 272K input tokens are billed at 2x the input/cache rates and 1.5x the output rate for the entire request, not just the excess; cache writes are 1.25x the uncached input rate.
Claude Fable 5.1Anthropic1.0M128K$10.00$50.00--Cache reads cut to $0.25/MTok, a 75% reduction from Fable 5's $1.00/MTok. 5-minute cache writes $12.50/MTok, 1-hour cache writes $20/MTok. Batch API is 50% off. Anthropic reports this brings typical-workload cost down about 25% and highly agentic workload cost down up to about 45% versus Fable 5 at the same prices.
Claude Mythos 5.1Anthropic1.0M128K$10.00$50.00--Identical pricing to Claude Fable 5.1: cache reads $0.25/MTok (2.5% of base input), 5-minute cache writes $12.50/MTok, 1-hour cache writes $20/MTok, Batch API 50% off.
Claude Opus 5.5Anthropic1.0M128K$4.00$20.00--Cache reads $0.20/MTok (5% of input), 5-min cache writes $5, 1-hour $8, Batch API 50% off. Source: platform.claude.com/docs/en/about-claude/pricing, checked 2026-09-26.
Claude Opus 5Anthropic1.0M--$5.00$25.0052.3
Claude Fable 5Anthropic1.0M128K$10.00$50.0070
Claude Sonnet 5Anthropic1.0M128K$2.00$10.0079Launch intro price of $2/$10 is now the standard price; the planned Sep 1 2026 rise to $3/$15 will not occur. Source: platform.claude.com/docs/en/about-claude/pricing, checked 2026-09-26.
Claude Opus 4.8Anthropic1.0M128K$5.00$25.0065
Claude Sonnet 4.6Anthropic1.0M128K$3.00$15.0047
GPT-5.5OpenAI1.0M--$5.00$30.0089
GPT-5.4OpenAI1.0M128K$2.50$15.00165
Gemini 3.5 FlashGoogle1.0M66K$1.50$9.00198
Gemini 3.8 FlashGoogle1.0M66K$0.75$3.75--$0.75/$3.75 through Dec 31 2026, then $1.50/$7.50 from Jan 1 2027 (same schedule as Gemini 3.7 Flash). Source: ai.google.dev/gemini-api/docs/pricing, checked 2026-09-26.
Gemini 2.5 ProGoogle1.0M--$1.25$10.00151$2.50/$15 per MTok for prompts over 200K tokens
Gemini 2.5 FlashGoogle1.0M66K$0.30$2.50222
GLM-5.2Z.ai1.0M131K$0.90$2.86--Via OpenRouter (model ID z-ai/glm-5.2); SiliconFlow lists $1.40/$4.40 per MTok
Llama 4 MaverickMeta1.0M--$0.15$0.60118Via third-party inference providers; Meta does not operate a paid API
DeepSeek V4 ProDeepSeek1.0M384K$0.43$0.877275% launch promo made permanent 2026-05-22 (original list was $1.74/$3.48 per MTok); $0.435 is now the standard rate.
DeepSeek V4 FlashDeepSeek1.0M384K$0.09$0.18113Cached input $0.0028/MTok
Qwen3.7 MaxAlibaba1.0M--$1.25$3.75206
Sakana Fugu UltraSakana AI1.0M131K$5.00$30.00--$10/$45 per MTok for contexts over 272K tokens; cached input $0.50/MTok (standard), $1.00/MTok (over 272K)
GPT-5OpenAI400K--$0.63$5.00107Succeeded by GPT-5.5 (official OpenAI replacement); no longer on pricing page. API shutdown Dec 11 2026.
Qwen3 MaxAlibaba262K--$0.78$3.9066
Tencent Hy3Tencent262K--$0.20$0.80--Free via OpenRouter for ~2 weeks from 2026-07-07 (model ID tencent/hy3:free); paid tier $0.20/$0.80 per MTok via OpenRouter
Claude Haiku 4.5Anthropic200K--$1.00$5.0098
o4-miniOpenAI200K--$1.10$4.40163
o3OpenAI200K--$2.00$8.00167
Grok 3xAI131K--$3.00$15.00--
DeepSeek R1DeepSeek128K--$0.55$2.19--deepseek-reasoner API alias now routes to V4 Flash thinking mode
Mistral Large 2Mistral128K--$2.00$6.0055
Mistral Small 3Mistral128K--$0.10$0.30166

Benchmark glossary

What each benchmark actually measures and why it matters.

SWE-bench Verified

Real GitHub issues from popular Python repos. The model must write code that passes the existing test suite. "Verified" means a human checked that each issue is solvable. Scores range 0-100%.

swebench.com

MMLU

Massive Multitask Language Understanding. 57 academic subjects from high-school to professional level (law, medicine, STEM). Tests breadth of knowledge. Scores are % correct.

Hendrycks et al.

HumanEval

OpenAI benchmark: 164 hand-crafted Python programming problems. Measures functional code generation (pass@1). Widely used but considered dated for frontier models.

OpenAI

MMMU

Massive Multidisciplinary Multimodal Understanding. 11.5K questions requiring images (charts, diagrams, photos) plus text. Standard for vision-language model evaluation.

mmmu-benchmark.github.io

RULER

Ruler for Long-Context Evaluation. Tests retrieval, multi-hop reasoning, and aggregation over very long documents (up to 128K tokens). Better signal than "needle in haystack" for long-context claims.

Hsieh et al.

GPQA Diamond

Graduate-Level Google-Proof Q&A, Diamond split. Expert-level science questions that Google cannot directly answer. High signal for deep scientific reasoning. Human experts score ~65%.

Rein et al.

MATH

Competition mathematics (AMC, AIME, Olympiad). 12,500 problems across 7 difficulty levels. Tests symbolic reasoning and multi-step derivation. Scores are % correct.

Hendrycks et al.

Not sure which model fits your project?

We evaluate models against your actual use case, run cost-to-quality tests across providers, and build the integration. No vendor lock-in, no guesswork.

Frequently asked questions

What is an LLM leaderboard?

An LLM leaderboard ranks AI language models by benchmark scores, speed, pricing, and context window so developers can pick the best model for their use case without reading every provider's documentation separately.

Which AI model scores highest on benchmarks in 2026?

Claude Fable 5 leads on SWE-bench Verified (95.0%) for coding. GPT-5-5 and Gemini 2.5 Pro are competitive on general reasoning. Use the filter above to sort by the benchmark most relevant to your workload.

How do I compare AI models side by side?

Use the leaderboard filters on this page, or go to the Compare tool to select up to four models and view their specs and benchmarks in a single table.

Which AI model API is the cheapest?

Open-weight models like Llama 4 Scout and Qwen3 are available free through several providers. Among paid APIs, DeepSeek V4 Pro offers strong reasoning at a lower cost per token than GPT-5 or Claude Fable 5. See the full pricing breakdown at /models/pricing.

Related guides

How we verify. Benchmark scores are taken from official provider pages, third-party leaderboards, or peer-reviewed papers. Where a number could not be independently confirmed it is shown as a dash. Pricing reflects the public API rate at the time of last update; check the provider for current pricing.

Affiliate disclosure. AY Automate has no affiliate relationship with any model provider listed here. Rankings are editorial, not commercial.