LLM API Pricing Comparison 2026
Set your monthly token volume and see real AI model costs across 36 models with published pricing. All prices are from provider documentation, not estimates. Batch discounts and prompt caching can cut costs by 50 to 90 percent; see the notes below.
Usage settings
Cheapest option
$1.26/mo
DeepSeek V4 Flash
Cost per request
$0.0013
at 1,000 req/mo (DeepSeek V4 Flash)
Price spread
158.7x
DeepSeek V4 Flash vs GPT-6 Astra
Monthly cost by model (cheapest first)
Showing top 15 of 36 models with published pricing. Hover for input/output breakdown.
| Model | Provider | Input cost | Output cost | Monthly total | Per request |
|---|---|---|---|---|---|
| DeepSeek | $0.90 | $0.36 | $1.26 | $0.0013 | |
| Meta | $0.80 | $0.60 | $1.40 | $0.0014 | |
| Mistral | $1.00 | $0.60 | $1.60 | $0.0016 | |
| OpenAI | $1.00 | $1.00 | $2.00 | $0.0020 | |
| Meta | $1.50 | $1.20 | $2.70 | $0.0027 | |
| Tencent | $2.00 | $1.60 | $3.60 | $0.0036 | |
| DeepSeek | $4.35 | $1.74 | $6.09 | $0.0061 | |
| $3.00 | $5.00 | $8.00 | $0.0080 | ||
| DeepSeek | $5.50 | $4.38 | $9.88 | $0.0099 | |
| Z.ai | $9.00 | $5.72 | $15 | $0.015 | |
| $7.50 | $7.50 | $15 | $0.015 | ||
| Alibaba | $7.80 | $7.80 | $16 | $0.016 | |
| OpenAI | $6.25 | $10 | $16 | $0.016 | |
| OpenAI | $11 | $8.80 | $20 | $0.020 | |
| Anthropic | $10 | $10 | $20 | $0.020 | |
| Alibaba | $13 | $7.50 | $20 | $0.020 | |
| Mistral | $20 | $12 | $32 | $0.032 | |
| xAI | $20 | $12 | $32 | $0.032 | |
| $13 | $20 | $33 | $0.033 | ||
| $15 | $18 | $33 | $0.033 | ||
| OpenAI | $20 | $16 | $36 | $0.036 | |
| Anthropic | $20 | $20 | $40 | $0.040 | |
| OpenAI | $20 | $20 | $40 | $0.040 | |
| OpenAI | $25 | $30 | $55 | $0.055 | |
| Anthropic | $30 | $30 | $60 | $0.060 | |
| xAI | $30 | $30 | $60 | $0.060 | |
| Anthropic | $40 | $40 | $80 | $0.080 | |
| OpenAI | $40 | $40 | $80 | $0.080 | |
| Anthropic | $50 | $50 | $100 | $0.100 | |
| Anthropic | $50 | $50 | $100 | $0.100 | |
| OpenAI | $50 | $60 | $110 | $0.110 | |
| Sakana AI | $50 | $60 | $110 | $0.110 | |
| Anthropic | $100 | $100 | $200 | $0.200 | |
| Anthropic | $100 | $100 | $200 | $0.200 | |
| Anthropic | $100 | $100 | $200 | $0.200 | |
| OpenAI | $100 | $100 | $200 | $0.200 |
Published pricing
Per-million-token rates from provider documentation. "Not published" means no standard public rate is listed. Sorted by input price, cheapest first.
| Model | Provider | In / MTok | Out / MTok | Note | Free tier |
|---|---|---|---|---|---|
| Llama 4 Scout | Meta | $0.08 | $0.30 | Via third-party inference providers; Meta does not operate a paid API | |
| DeepSeek V4 Flash | DeepSeek | $0.09 | $0.18 | Cached input $0.0028/MTok | |
| GPT-6 Luna | OpenAI | $0.10 | $0.50 | Cached input $0.01/MTok. Prompts over 272K input tokens cost 2x input and cache rates and 1.5x output. Source: developers.openai.com/api/docs/models/gpt-6-luna, checked 2026-09-26. | |
| Mistral Small 3 | Mistral | $0.10 | $0.30 | ||
| Llama 4 Maverick | Meta | $0.15 | $0.60 | Via third-party inference providers; Meta does not operate a paid API | |
| Tencent Hy3 | Tencent | $0.20 | $0.80 | Free via OpenRouter for ~2 weeks from 2026-07-07 (model ID tencent/hy3:free); paid tier $0.20/$0.80 per MTok via OpenRouter | Free tier |
| Gemini 2.5 Flash | $0.30 | $2.50 | |||
| DeepSeek V4 Pro | DeepSeek | $0.43 | $0.87 | 75% launch promo made permanent 2026-05-22 (original list was $1.74/$3.48 per MTok); $0.435 is now the standard rate. | |
| DeepSeek R1 | DeepSeek | $0.55 | $2.19 | deepseek-reasoner API alias now routes to V4 Flash thinking mode | |
| GPT-5 | OpenAI | $0.63 | $5.00 | Succeeded by GPT-5.5 (official OpenAI replacement); no longer on pricing page. API shutdown Dec 11 2026. | |
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.75/$3.75 through Dec 31 2026, then $1.50/$7.50 from Jan 1 2027 (same schedule as Gemini 3.7 Flash). Source: ai.google.dev/gemini-api/docs/pricing, checked 2026-09-26. | ||
| Qwen3 Max | Alibaba | $0.78 | $3.90 | ||
| GLM-5.2 | Z.ai | $0.90 | $2.86 | Via OpenRouter (model ID z-ai/glm-5.2); SiliconFlow lists $1.40/$4.40 per MTok | |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | ||
| o4-mini | OpenAI | $1.10 | $4.40 | ||
| Gemini 2.5 Pro | $1.25 | $10.00 | $2.50/$15 per MTok for prompts over 200K tokens | ||
| Qwen3.7 Max | Alibaba | $1.25 | $3.75 | ||
| Gemini 3.5 Flash | $1.50 | $9.00 | |||
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | Launch intro price of $2/$10 is now the standard price; the planned Sep 1 2026 rise to $3/$15 will not occur. Source: platform.claude.com/docs/en/about-claude/pricing, checked 2026-09-26. | |
| GPT-6 Sol | OpenAI | $2.00 | $10.00 | Cached input $0.20/MTok. Prompts over 272K input tokens cost 2x input and cache rates and 1.5x output for the full request. Source: developers.openai.com/api/docs/models/gpt-6-sol, checked 2026-09-26. | |
| o3 | OpenAI | $2.00 | $8.00 | ||
| Mistral Large 2 | Mistral | $2.00 | $6.00 | ||
| Grok 4.20 | xAI | $2.00 | $6.00 | ||
| GPT-5.4 | OpenAI | $2.50 | $15.00 | ||
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | ||
| Grok 3 | xAI | $3.00 | $15.00 | ||
| Claude Opus 5.5 | Anthropic | $4.00 | $20.00 | Cache reads $0.20/MTok (5% of input), 5-min cache writes $5, 1-hour $8, Batch API 50% off. Source: platform.claude.com/docs/en/about-claude/pricing, checked 2026-09-26. | |
| GPT-5.6 Sol | OpenAI | $4.00 | $20.00 | OpenAI cut pricing from launch $5/$30 to $4/$20 per MTok on Aug 21 2026, a promotion running at least through Nov 21 2026. Prompts over 272K input tokens are billed at 2x the input/cache rates and 1.5x the output rate for the entire request, not just the excess; cache writes are 1.25x the uncached input rate. | |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | ||
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | ||
| GPT-5.5 | OpenAI | $5.00 | $30.00 | ||
| Sakana Fugu Ultra | Sakana AI | $5.00 | $30.00 | $10/$45 per MTok for contexts over 272K tokens; cached input $0.50/MTok (standard), $1.00/MTok (over 272K) | |
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | Cache reads cut to $0.25/MTok, a 75% reduction from Fable 5's $1.00/MTok. 5-minute cache writes $12.50/MTok, 1-hour cache writes $20/MTok. Batch API is 50% off. Anthropic reports this brings typical-workload cost down about 25% and highly agentic workload cost down up to about 45% versus Fable 5 at the same prices. | |
| Claude Mythos 5.1 | Anthropic | $10.00 | $50.00 | Identical pricing to Claude Fable 5.1: cache reads $0.25/MTok (2.5% of base input), 5-minute cache writes $12.50/MTok, 1-hour cache writes $20/MTok, Batch API 50% off. | |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | ||
| GPT-6 Astra | OpenAI | $10.00 | $50.00 | Cached input $1/MTok, cache writes $12.50/MTok. Prompts over 272K input tokens are billed at 2x the input/cache rates and 1.5x the output rate for the entire request, not just the excess. | |
| Gemini 3.5 Pro | Not published | Not published | GA expected Jul 17 2026; pricing not confirmed as of 2026-07-06 |
Cost to fill the context window
For long-document and RAG workloads, the input price per million tokens does not tell the full story. A cheaper model with a smaller context window can cost more per useful context than a pricier model with a large window. Sorted cheapest first.
| Model | Provider | Context window | In / MTok | Cost to fill context |
|---|---|---|---|---|
| Mistral Small 3 | Mistral | 128K | $0.10 | $0.013 |
| Tencent Hy3 | Tencent | 262K | $0.20 | $0.052 |
| DeepSeek R1 | DeepSeek | 128K | $0.55 | $0.070 |
| DeepSeek V4 Flash | DeepSeek | 1.0M | $0.09 | $0.090 |
| GPT-6 Luna | OpenAI | 1.1M | $0.10 | $0.105 |
| Llama 4 Maverick | Meta | 1.0M | $0.15 | $0.150 |
| Claude Haiku 4.5 | Anthropic | 200K | $1.00 | $0.200 |
| Qwen3 Max | Alibaba | 262K | $0.78 | $0.204 |
| o4-mini | OpenAI | 200K | $1.10 | $0.220 |
| GPT-5 | OpenAI | 400K | $0.63 | $0.250 |
| Mistral Large 2 | Mistral | 128K | $2.00 | $0.256 |
| Gemini 2.5 Flash | 1.0M | $0.30 | $0.315 | |
| Grok 3 | xAI | 131K | $3.00 | $0.393 |
| o3 | OpenAI | 200K | $2.00 | $0.400 |
| DeepSeek V4 Pro | DeepSeek | 1.0M | $0.43 | $0.435 |
| Gemini 3.8 Flash | 1.0M | $0.75 | $0.786 | |
| Llama 4 Scout | Meta | 10.0M | $0.08 | $0.800 |
| GLM-5.2 | Z.ai | 1.0M | $0.90 | $0.944 |
| Qwen3.7 Max | Alibaba | 1.0M | $1.25 | $1.25 |
| Gemini 2.5 Pro | 1.0M | $1.25 | $1.31 | |
| Gemini 3.5 Flash | 1.0M | $1.50 | $1.57 | |
| Claude Sonnet 5 | Anthropic | 1.0M | $2.00 | $2.10 |
| GPT-6 Sol | OpenAI | 1.1M | $2.00 | $2.10 |
| GPT-5.4 | OpenAI | 1.0M | $2.50 | $2.62 |
| Claude Sonnet 4.6 | Anthropic | 1.0M | $3.00 | $3.15 |
| Grok 4.20 | xAI | 2.0M | $2.00 | $4.00 |
| Claude Opus 5.5 | Anthropic | 1.0M | $4.00 | $4.19 |
| GPT-5.6 Sol | OpenAI | 1.1M | $4.00 | $4.20 |
| Sakana Fugu Ultra | Sakana AI | 1.0M | $5.00 | $5.00 |
| Claude Opus 5 | Anthropic | 1.0M | $5.00 | $5.24 |
| Claude Opus 4.8 | Anthropic | 1.0M | $5.00 | $5.24 |
| GPT-5.5 | OpenAI | 1.0M | $5.00 | $5.24 |
| Claude Fable 5.1 | Anthropic | 1.0M | $10.00 | $10.49 |
| Claude Mythos 5.1 | Anthropic | 1.0M | $10.00 | $10.49 |
| Claude Fable 5 | Anthropic | 1.0M | $10.00 | $10.49 |
| GPT-6 Astra | OpenAI | 1.1M | $10.00 | $10.50 |
What the calculator does not include
- Batch processing (50% off). Anthropic, OpenAI, and Google all offer asynchronous batch APIs at half the standard rate. Good for offline jobs: embeddings, classification, bulk summarization. See anthropic.com/pricing, openai.com/api/pricing, and ai.google.dev/pricing.
- Prompt caching (up to 90% off cached input). Claude and Gemini charge as little as 10 percent of the standard input rate for cache hits. If your workload reuses a long system prompt or repeated document, real costs can be 80 to 90 percent lower than the calculator shows.
- Volume and enterprise discounts. At sustained high volumes, providers negotiate rates not reflected in public pricing.
- Free tiers. Several providers offer free rate-limited access. See the free models directory.
- Self-hosting open-weight models. Llama, Qwen, Mistral, and DeepSeek R1 can be self-hosted. You pay infrastructure, not per-token rates. At sustained scale, self-hosting can be meaningfully cheaper than API pricing for output-heavy workloads.
Not sure which model fits your budget and quality bar?
We run cost-to-quality evaluations against your actual prompts and data, benchmark multiple providers, and build the integration. No vendor lock-in, no guesswork.
Related
LLM pricing: common questions
Which LLM API is cheapest in 2026?
By input price per 1M tokens, the cheapest models in the table above are Llama 4 Scout ($0.08 in / $0.30 out), DeepSeek V4 Flash ($0.09 in / $0.18 out), GPT-6 Luna ($0.10 in / $0.50 out). The cheapest closed-weight API is GPT-6 Luna at $0.10 input. The lowest output rate is DeepSeek V4 Flash at $0.18. Open-weight model prices come from third-party hosts and vary by provider. A low per-token rate is not always the cheapest task: check output price and how many tokens the model needs.
Do any AI models offer a free API tier?
Yes. Llama 4, Gemini 2.5 Flash, and Qwen3 all offer free rate-limited access. See the free models directory for the full list with usage limits.
How much does batch API pricing save?
Anthropic, OpenAI, and Google all offer asynchronous batch APIs at 50% off standard rates. Suitable for offline jobs: embeddings, classification, bulk summarization.
How much does prompt caching reduce costs?
Cache hits on most Claude models cost 10% of the standard input rate, 5% on Claude Opus 5.5 and 2.5% on Claude Fable 5.1, per Anthropic's pricing page. OpenAI lists cached input at 10% of the input rate for GPT-6 Astra, Sol and Luna. For workloads with long repeated system prompts or documents, real costs can be 80 to 90% lower than the calculator's base estimate.
Pricing accuracy. All prices reflect published public API rates at the time of last update. Prices change. Verify at the provider before committing to a budget. Models without a published rate are excluded from the calculator.
Affiliate disclosure. AY Automate has no affiliate relationship with any model provider listed here. Rankings and recommendations are editorial, not commercial.