Pricing

LLM API Pricing Comparison 2026

Set your monthly token volume and see real AI model costs across 36 models with published pricing. All prices are from provider documentation, not estimates. Batch discounts and prompt caching can cut costs by 50 to 90 percent; see the notes below.

Usage settings

M tokens
M tokens

Cheapest option

$1.26/mo

DeepSeek V4 Flash

Cost per request

$0.0013

at 1,000 req/mo (DeepSeek V4 Flash)

Price spread

158.7x

DeepSeek V4 Flash vs GPT-6 Astra

Monthly cost by model (cheapest first)

Showing top 15 of 36 models with published pricing. Hover for input/output breakdown.

ModelProviderInput costOutput costMonthly totalPer request
DeepSeek V4 FlashCached input $0.0028/MTokDeepSeek$0.90$0.36$1.26$0.0013
Llama 4 ScoutVia third-party inference providers; Meta does not operate a paid APIMeta$0.80$0.60$1.40$0.0014
Mistral Small 3Mistral$1.00$0.60$1.60$0.0016
GPT-6 LunaCached input $0.01/MTok. Prompts over 272K input tokens cost 2x input and cache rates and 1.5x output. Source: developers.openai.com/api/docs/models/gpt-6-luna, checked 2026-09-26.OpenAI$1.00$1.00$2.00$0.0020
Llama 4 MaverickVia third-party inference providers; Meta does not operate a paid APIMeta$1.50$1.20$2.70$0.0027
Tencent Hy3Free via OpenRouter for ~2 weeks from 2026-07-07 (model ID tencent/hy3:free); paid tier $0.20/$0.80 per MTok via OpenRouterTencent$2.00$1.60$3.60$0.0036
DeepSeek V4 Pro75% launch promo made permanent 2026-05-22 (original list was $1.74/$3.48 per MTok); $0.435 is now the standard rate.DeepSeek$4.35$1.74$6.09$0.0061
Gemini 2.5 FlashGoogle$3.00$5.00$8.00$0.0080
DeepSeek R1deepseek-reasoner API alias now routes to V4 Flash thinking modeDeepSeek$5.50$4.38$9.88$0.0099
GLM-5.2Via OpenRouter (model ID z-ai/glm-5.2); SiliconFlow lists $1.40/$4.40 per MTokZ.ai$9.00$5.72$15$0.015
Gemini 3.8 Flash$0.75/$3.75 through Dec 31 2026, then $1.50/$7.50 from Jan 1 2027 (same schedule as Gemini 3.7 Flash). Source: ai.google.dev/gemini-api/docs/pricing, checked 2026-09-26.Google$7.50$7.50$15$0.015
Qwen3 MaxAlibaba$7.80$7.80$16$0.016
GPT-5Succeeded by GPT-5.5 (official OpenAI replacement); no longer on pricing page. API shutdown Dec 11 2026.OpenAI$6.25$10$16$0.016
o4-miniOpenAI$11$8.80$20$0.020
Claude Haiku 4.5Anthropic$10$10$20$0.020
Qwen3.7 MaxAlibaba$13$7.50$20$0.020
Mistral Large 2Mistral$20$12$32$0.032
Grok 4.20xAI$20$12$32$0.032
Gemini 2.5 Pro$2.50/$15 per MTok for prompts over 200K tokensGoogle$13$20$33$0.033
Gemini 3.5 FlashGoogle$15$18$33$0.033
o3OpenAI$20$16$36$0.036
Claude Sonnet 5Launch intro price of $2/$10 is now the standard price; the planned Sep 1 2026 rise to $3/$15 will not occur. Source: platform.claude.com/docs/en/about-claude/pricing, checked 2026-09-26.Anthropic$20$20$40$0.040
GPT-6 SolCached input $0.20/MTok. Prompts over 272K input tokens cost 2x input and cache rates and 1.5x output for the full request. Source: developers.openai.com/api/docs/models/gpt-6-sol, checked 2026-09-26.OpenAI$20$20$40$0.040
GPT-5.4OpenAI$25$30$55$0.055
Claude Sonnet 4.6Anthropic$30$30$60$0.060
Grok 3xAI$30$30$60$0.060
Claude Opus 5.5Cache reads $0.20/MTok (5% of input), 5-min cache writes $5, 1-hour $8, Batch API 50% off. Source: platform.claude.com/docs/en/about-claude/pricing, checked 2026-09-26.Anthropic$40$40$80$0.080
GPT-5.6 SolOpenAI cut pricing from launch $5/$30 to $4/$20 per MTok on Aug 21 2026, a promotion running at least through Nov 21 2026. Prompts over 272K input tokens are billed at 2x the input/cache rates and 1.5x the output rate for the entire request, not just the excess; cache writes are 1.25x the uncached input rate.OpenAI$40$40$80$0.080
Claude Opus 5Anthropic$50$50$100$0.100
Claude Opus 4.8Anthropic$50$50$100$0.100
GPT-5.5OpenAI$50$60$110$0.110
Sakana Fugu Ultra$10/$45 per MTok for contexts over 272K tokens; cached input $0.50/MTok (standard), $1.00/MTok (over 272K)Sakana AI$50$60$110$0.110
Claude Fable 5.1Cache reads cut to $0.25/MTok, a 75% reduction from Fable 5's $1.00/MTok. 5-minute cache writes $12.50/MTok, 1-hour cache writes $20/MTok. Batch API is 50% off. Anthropic reports this brings typical-workload cost down about 25% and highly agentic workload cost down up to about 45% versus Fable 5 at the same prices.Anthropic$100$100$200$0.200
Claude Mythos 5.1Identical pricing to Claude Fable 5.1: cache reads $0.25/MTok (2.5% of base input), 5-minute cache writes $12.50/MTok, 1-hour cache writes $20/MTok, Batch API 50% off.Anthropic$100$100$200$0.200
Claude Fable 5Anthropic$100$100$200$0.200
GPT-6 AstraCached input $1/MTok, cache writes $12.50/MTok. Prompts over 272K input tokens are billed at 2x the input/cache rates and 1.5x the output rate for the entire request, not just the excess.OpenAI$100$100$200$0.200

Published pricing

Per-million-token rates from provider documentation. "Not published" means no standard public rate is listed. Sorted by input price, cheapest first.

ModelProviderIn / MTokOut / MTokNoteFree tier
Llama 4 ScoutMeta$0.08$0.30Via third-party inference providers; Meta does not operate a paid API
DeepSeek V4 FlashDeepSeek$0.09$0.18Cached input $0.0028/MTok
GPT-6 LunaOpenAI$0.10$0.50Cached input $0.01/MTok. Prompts over 272K input tokens cost 2x input and cache rates and 1.5x output. Source: developers.openai.com/api/docs/models/gpt-6-luna, checked 2026-09-26.
Mistral Small 3Mistral$0.10$0.30
Llama 4 MaverickMeta$0.15$0.60Via third-party inference providers; Meta does not operate a paid API
Tencent Hy3Tencent$0.20$0.80Free via OpenRouter for ~2 weeks from 2026-07-07 (model ID tencent/hy3:free); paid tier $0.20/$0.80 per MTok via OpenRouterFree tier
Gemini 2.5 FlashGoogle$0.30$2.50
DeepSeek V4 ProDeepSeek$0.43$0.8775% launch promo made permanent 2026-05-22 (original list was $1.74/$3.48 per MTok); $0.435 is now the standard rate.
DeepSeek R1DeepSeek$0.55$2.19deepseek-reasoner API alias now routes to V4 Flash thinking mode
GPT-5OpenAI$0.63$5.00Succeeded by GPT-5.5 (official OpenAI replacement); no longer on pricing page. API shutdown Dec 11 2026.
Gemini 3.8 FlashGoogle$0.75$3.75$0.75/$3.75 through Dec 31 2026, then $1.50/$7.50 from Jan 1 2027 (same schedule as Gemini 3.7 Flash). Source: ai.google.dev/gemini-api/docs/pricing, checked 2026-09-26.
Qwen3 MaxAlibaba$0.78$3.90
GLM-5.2Z.ai$0.90$2.86Via OpenRouter (model ID z-ai/glm-5.2); SiliconFlow lists $1.40/$4.40 per MTok
Claude Haiku 4.5Anthropic$1.00$5.00
o4-miniOpenAI$1.10$4.40
Gemini 2.5 ProGoogle$1.25$10.00$2.50/$15 per MTok for prompts over 200K tokens
Qwen3.7 MaxAlibaba$1.25$3.75
Gemini 3.5 FlashGoogle$1.50$9.00
Claude Sonnet 5Anthropic$2.00$10.00Launch intro price of $2/$10 is now the standard price; the planned Sep 1 2026 rise to $3/$15 will not occur. Source: platform.claude.com/docs/en/about-claude/pricing, checked 2026-09-26.
GPT-6 SolOpenAI$2.00$10.00Cached input $0.20/MTok. Prompts over 272K input tokens cost 2x input and cache rates and 1.5x output for the full request. Source: developers.openai.com/api/docs/models/gpt-6-sol, checked 2026-09-26.
o3OpenAI$2.00$8.00
Mistral Large 2Mistral$2.00$6.00
Grok 4.20xAI$2.00$6.00
GPT-5.4OpenAI$2.50$15.00
Claude Sonnet 4.6Anthropic$3.00$15.00
Grok 3xAI$3.00$15.00
Claude Opus 5.5Anthropic$4.00$20.00Cache reads $0.20/MTok (5% of input), 5-min cache writes $5, 1-hour $8, Batch API 50% off. Source: platform.claude.com/docs/en/about-claude/pricing, checked 2026-09-26.
GPT-5.6 SolOpenAI$4.00$20.00OpenAI cut pricing from launch $5/$30 to $4/$20 per MTok on Aug 21 2026, a promotion running at least through Nov 21 2026. Prompts over 272K input tokens are billed at 2x the input/cache rates and 1.5x the output rate for the entire request, not just the excess; cache writes are 1.25x the uncached input rate.
Claude Opus 5Anthropic$5.00$25.00
Claude Opus 4.8Anthropic$5.00$25.00
GPT-5.5OpenAI$5.00$30.00
Sakana Fugu UltraSakana AI$5.00$30.00$10/$45 per MTok for contexts over 272K tokens; cached input $0.50/MTok (standard), $1.00/MTok (over 272K)
Claude Fable 5.1Anthropic$10.00$50.00Cache reads cut to $0.25/MTok, a 75% reduction from Fable 5's $1.00/MTok. 5-minute cache writes $12.50/MTok, 1-hour cache writes $20/MTok. Batch API is 50% off. Anthropic reports this brings typical-workload cost down about 25% and highly agentic workload cost down up to about 45% versus Fable 5 at the same prices.
Claude Mythos 5.1Anthropic$10.00$50.00Identical pricing to Claude Fable 5.1: cache reads $0.25/MTok (2.5% of base input), 5-minute cache writes $12.50/MTok, 1-hour cache writes $20/MTok, Batch API 50% off.
Claude Fable 5Anthropic$10.00$50.00
GPT-6 AstraOpenAI$10.00$50.00Cached input $1/MTok, cache writes $12.50/MTok. Prompts over 272K input tokens are billed at 2x the input/cache rates and 1.5x the output rate for the entire request, not just the excess.
Gemini 3.5 ProGoogleNot publishedNot publishedGA expected Jul 17 2026; pricing not confirmed as of 2026-07-06

Cost to fill the context window

For long-document and RAG workloads, the input price per million tokens does not tell the full story. A cheaper model with a smaller context window can cost more per useful context than a pricier model with a large window. Sorted cheapest first.

ModelProviderContext windowIn / MTokCost to fill context
Mistral Small 3Mistral128K$0.10$0.013
Tencent Hy3Tencent262K$0.20$0.052
DeepSeek R1DeepSeek128K$0.55$0.070
DeepSeek V4 FlashDeepSeek1.0M$0.09$0.090
GPT-6 LunaOpenAI1.1M$0.10$0.105
Llama 4 MaverickMeta1.0M$0.15$0.150
Claude Haiku 4.5Anthropic200K$1.00$0.200
Qwen3 MaxAlibaba262K$0.78$0.204
o4-miniOpenAI200K$1.10$0.220
GPT-5OpenAI400K$0.63$0.250
Mistral Large 2Mistral128K$2.00$0.256
Gemini 2.5 FlashGoogle1.0M$0.30$0.315
Grok 3xAI131K$3.00$0.393
o3OpenAI200K$2.00$0.400
DeepSeek V4 ProDeepSeek1.0M$0.43$0.435
Gemini 3.8 FlashGoogle1.0M$0.75$0.786
Llama 4 ScoutMeta10.0M$0.08$0.800
GLM-5.2Z.ai1.0M$0.90$0.944
Qwen3.7 MaxAlibaba1.0M$1.25$1.25
Gemini 2.5 ProGoogle1.0M$1.25$1.31
Gemini 3.5 FlashGoogle1.0M$1.50$1.57
Claude Sonnet 5Anthropic1.0M$2.00$2.10
GPT-6 SolOpenAI1.1M$2.00$2.10
GPT-5.4OpenAI1.0M$2.50$2.62
Claude Sonnet 4.6Anthropic1.0M$3.00$3.15
Grok 4.20xAI2.0M$2.00$4.00
Claude Opus 5.5Anthropic1.0M$4.00$4.19
GPT-5.6 SolOpenAI1.1M$4.00$4.20
Sakana Fugu UltraSakana AI1.0M$5.00$5.00
Claude Opus 5Anthropic1.0M$5.00$5.24
Claude Opus 4.8Anthropic1.0M$5.00$5.24
GPT-5.5OpenAI1.0M$5.00$5.24
Claude Fable 5.1Anthropic1.0M$10.00$10.49
Claude Mythos 5.1Anthropic1.0M$10.00$10.49
Claude Fable 5Anthropic1.0M$10.00$10.49
GPT-6 AstraOpenAI1.1M$10.00$10.50

What the calculator does not include

  • Batch processing (50% off). Anthropic, OpenAI, and Google all offer asynchronous batch APIs at half the standard rate. Good for offline jobs: embeddings, classification, bulk summarization. See anthropic.com/pricing, openai.com/api/pricing, and ai.google.dev/pricing.
  • Prompt caching (up to 90% off cached input). Claude and Gemini charge as little as 10 percent of the standard input rate for cache hits. If your workload reuses a long system prompt or repeated document, real costs can be 80 to 90 percent lower than the calculator shows.
  • Volume and enterprise discounts. At sustained high volumes, providers negotiate rates not reflected in public pricing.
  • Free tiers. Several providers offer free rate-limited access. See the free models directory.
  • Self-hosting open-weight models. Llama, Qwen, Mistral, and DeepSeek R1 can be self-hosted. You pay infrastructure, not per-token rates. At sustained scale, self-hosting can be meaningfully cheaper than API pricing for output-heavy workloads.

Not sure which model fits your budget and quality bar?

We run cost-to-quality evaluations against your actual prompts and data, benchmark multiple providers, and build the integration. No vendor lock-in, no guesswork.

Related

LLM pricing: common questions

Which LLM API is cheapest in 2026?

By input price per 1M tokens, the cheapest models in the table above are Llama 4 Scout ($0.08 in / $0.30 out), DeepSeek V4 Flash ($0.09 in / $0.18 out), GPT-6 Luna ($0.10 in / $0.50 out). The cheapest closed-weight API is GPT-6 Luna at $0.10 input. The lowest output rate is DeepSeek V4 Flash at $0.18. Open-weight model prices come from third-party hosts and vary by provider. A low per-token rate is not always the cheapest task: check output price and how many tokens the model needs.

Do any AI models offer a free API tier?

Yes. Llama 4, Gemini 2.5 Flash, and Qwen3 all offer free rate-limited access. See the free models directory for the full list with usage limits.

How much does batch API pricing save?

Anthropic, OpenAI, and Google all offer asynchronous batch APIs at 50% off standard rates. Suitable for offline jobs: embeddings, classification, bulk summarization.

How much does prompt caching reduce costs?

Cache hits on most Claude models cost 10% of the standard input rate, 5% on Claude Opus 5.5 and 2.5% on Claude Fable 5.1, per Anthropic's pricing page. OpenAI lists cached input at 10% of the input rate for GPT-6 Astra, Sol and Luna. For workloads with long repeated system prompts or documents, real costs can be 80 to 90% lower than the calculator's base estimate.

Pricing accuracy. All prices reflect published public API rates at the time of last update. Prices change. Verify at the provider before committing to a budget. Models without a published rate are excluded from the calculator.

Affiliate disclosure. AY Automate has no affiliate relationship with any model provider listed here. Rankings and recommendations are editorial, not commercial.