LLM API Pricing Comparison - July 2026

Verified July 13: GPT-5.6 Sol/Terra/Luna hit GA at $5/$30, $2.50/$15 and $1/$6, Grok 4.5 lands at $2/$6, Grok 4.1 Fast quietly redirects to pricier Grok 4.3, DeepSeek's peak-hour doubling goes live July 15.

Cheapest: Ministral 3B Best Value: DeepSeek V4 Flash Updated weekly
LLM API Pricing Comparison - July 2026

TL;DR

  • GPT-5.6 Sol/Terra/Luna went GA July 9 - Sol matches GPT-5.5 at $5/$30, Terra matches the soon-retired GPT-5.4 at $2.50/$15, Luna opens a new $1/$6 tier
  • Grok 4.1 Fast is no longer the cheap long-context pick - the slug quietly redirects to Grok 4.3 pricing ($1.25/$2.50) since May 15 and fully retires August 15
  • Grok 4.5 launched at $2/$6 with a 500K context window, undercutting GPT-5.4 mini on output while beating it on quality claims

The Bottom Line

GPT-5.6 is the headline change this week. OpenAI previewed the three-tier family - Sol, Terra, Luna - to a small government-vetted group on June 26, then took it fully generally available on July 9 after clearing an additional round of review tied to the administration's AI cybersecurity order. Pricing lands exactly where the preview numbers pointed: Sol at $5/$30 (matching GPT-5.5), Terra at $2.50/$15 (matching GPT-5.4), and Luna as a new bargain tier at $1/$6. All three carry a 1.05M token context window, 128K max output, and cached input at 10% of the standard rate. GPT-5.4 itself retires July 23, so Terra is effectively its replacement at the same price.

The more consequential correction this week is on the cheap end of the xAI lineup. Grok 4.1 Fast has been listed in this table for months at $0.20/$0.50 with a 2M context window - and that pricing hasn't been real since May 15, 2026. xAI retired the slug that day along with seven other legacy models; requests still resolve without erroring, but they silently redirect to Grok 4.3 and bill at $1.25/$2.50, more than 6x the old input rate. The slug disappears completely on August 15. Anyone still routing to grok-4-1-fast-reasoning or -non-reasoning in production has been overpaying for two months without a code change to show for it - check your xAI invoices against the model field, not just the slug you sent.

Grok 4.5 is the actual new xAI model worth assessing. It launched publicly July 8 at $2/$6 with a 500K context window (down from Grok 4.3's 1M) and cached input at $0.50 - or $1.00 above 200K tokens. It's a 1.5T-parameter MoE model, xAI's first V9-architecture model at full scale, and token-efficient in practice: independent benchmarks put it well behind Opus 4.8 and Fable 5 on coding quality despite provider-run harnesses showing closer results.

DeepSeek's peak-hour pricing is no longer a future announcement - the official V4 release lands July 15, two days after this update. Off-peak V4 Flash stays at $0.14/$0.28; peak-hour (9AM-noon and 2-6PM Beijing time) rises to $0.28/$0.56. DeepSeek V4 Flash remains the best raw value for workloads that can run off-peak or stay in US timezones.

Full Pricing Table

All prices in USD per million tokens (MTok). Verified against official documentation July 13, 2026. Sorted by input price, cheapest first.

ModelProviderInput (/1M)Output (/1M)ContextNotes
Ministral 3BMistral$0.04$0.04256KLegacy endpoint only
Llama 3.1 8BGroq$0.05$0.08128K840+ tok/s on LPU
GPT OSS 20BGroq$0.075$0.30128KOpen-weight on LPU
Mistral Small 3.2Mistral$0.08$0.20128KCheapest Mistral with function calling
GPT-4.1 nanoOpenAI$0.10$0.401MRouting and classification
Gemini 2.5 Flash-LiteGoogle$0.10$0.401MFree tier
Mistral Small 4Mistral$0.10$0.30128K
Ministral 3 3B 2512Mistral$0.10$0.10256KCurrent gen; legacy 3B endpoint at $0.04
Llama 4 ScoutGroq$0.11$0.34128K17B active / 16E MoE
DeepSeek V4 FlashDeepSeek$0.14$0.281MCache hit $0.0028; best value - peak-hour doubling goes official Jul 15
Ministral 3-8BMistral$0.15$0.15256K
GPT OSS 120BGroq$0.15$0.60128K
GPT-5.6 LunaOpenAI$0.20$1.251.05MNew Jul 9 GA; cached $0.02
GPT-5.4 nanoOpenAI$0.20$1.25n/aCached input $0.02
Ministral 3-14BMistral$0.20$0.20256KFlat in/out; optional reasoning mode
Gemini 3.1 Flash-LiteGoogle$0.25$1.501MFree tier
Qwen3 32BGroq$0.29$0.59131KOpen-weight; strong multilingual
Gemini 2.5 FlashGoogle$0.30$2.501MFree tier; solid mid-range
GPT-4.1 miniOpenAI$0.40$1.601M
Devstral 2Mistral$0.40$2.00256KCoding/dev agent specialized
DeepSeek V4 ProDeepSeek$0.435$0.871MCache hit $0.003625; peak-hour doubling goes official Jul 15
Mistral Large 3Mistral$0.50$1.50262K75% below Mistral Large 2
Llama 3.3 70BGroq$0.59$0.79128KDense 70B; instruction following
GPT-5.4 miniOpenAI$0.75$4.50400KCached $0.075; retires with GPT-5.4 Jul 23
Claude Haiku 4.5Anthropic$1.00$5.00200KCheapest active Anthropic model
Grok Build 0.1xAI$1.00$2.00256KAgentic coding; native MCP
o4-miniOpenAI$1.10$4.40200KCheapest dedicated reasoning model
Grok 4.3xAI$1.25$2.501MNow also the effective price for redirected Grok 4.1 Fast traffic
Grok 4.20xAI$1.25$2.501MReasoning variant; same price as 4.3
Gemini 2.5 ProGoogle$1.25$10.001M$2.50/$15 above 200K tokens
GPT-5.1OpenAI$1.25$10.00128K
GLM-5.2Z.AI$1.40$4.401MOpen weights; cache hit $0.26
Gemini 3.5 FlashGoogle$1.50$9.001MFree tier; batch $0.75/$4.50
Mistral Medium 3.5Mistral$1.50$7.50128K
GPT-5.2OpenAI$1.75$14.00128K
GPT-5.3-codexOpenAI$1.75$14.00128KCode specialist
o3OpenAI$2.00$8.00200K80% cut from original o1 pricing
GPT-4.1OpenAI$2.00$8.001M
Gemini 3.1 Pro PreviewGoogle$2.00$12.001M$4/$18 above 200K tokens
Claude Sonnet 5Anthropic$2.00$10.001MIntro through Aug 31; $3/$15 after; new tokenizer adds ~30% tokens
Grok 4.5xAI$2.00$6.00500KNew Jul 8; cached $0.50 ($1.00 above 200K)
GPT-5.6 TerraOpenAI$2.50$15.001.05MNew Jul 9 GA; matches retiring GPT-5.4 price
GPT-5.4OpenAI$2.50$15.001MRetires Jul 23; replaced by GPT-5.6 Terra
Claude Sonnet 4.6Anthropic$3.00$15.001M
Claude Opus 4.7Anthropic$5.00$25.001MFast Mode deprecated Jul 24
Claude Opus 4.8Anthropic$5.00$25.001MFast Mode $10/$50; standard Anthropic flagship
GPT-5.6 SolOpenAI$5.00$30.001.05MNew Jul 9 GA; matches GPT-5.5 price
GPT-5.5OpenAI$5.00$30.001MStays available alongside GPT-5.6
Claude Fable 5Anthropic$10.00$50.001MRestored Jul 1; suspended Jun 12-Jun 30
GPT-5.4 Pro / GPT-5.5 ProOpenAI$30.00$180.001MResearch/enterprise tier

Grok 4.1 Fast is omitted above because it no longer has its own price - see "Grok 4.1 Fast Silently Repriced" under Hidden Costs.

For benchmark context behind these prices, see the cost-efficiency leaderboard.

A calculator and pen on a notebook, representing the cost math behind API pricing decisions Working out real API costs requires more than reading the headline price - caching, batching, context surcharges, and tokenizer differences all shift the final number. Source: unsplash.com

Changes Since July 6

Four standout changes this week.

GPT-5.6 goes GA (July 9, 2026) - GPT-5.6 Sol, Terra, and Luna cleared the small-partner preview and shipped to everyone July 9, three days after this table was last verified. Pricing matches what leaked during preview: Sol $5/$30 (same as GPT-5.5), Terra $2.50/$15 (same as the outgoing GPT-5.4), Luna $1/$6 as a new budget tier with a full 1.05M context window. Cached input is 10% of standard across all three. GPT-5.4 itself retires July 23 and its traffic rolls onto Terra at the same price - a rare same-cost forced migration.

Grok 4.1 Fast quietly repriced (effective since May 15, 2026) - This is a correction, not new news: Grok 4.1 Fast has not actually billed at its listed $0.20/$0.50 rate since May 15. xAI retired the slug that day along with seven other legacy models and pointed all requests to Grok 4.3 pricing ($1.25/$2.50) instead of returning an error. We're pulling the stale row from this table now and flagging it under Hidden Costs below - anyone who copied last month's numbers into a cost model has been under-forecasting by more than 6x on input.

Grok 4.5 launches (July 8, 2026) - Grok 4.5 is xAI's first V9-architecture model at full scale, a 1.5T-parameter MoE priced at $2/$6 with a 500K context window (down from Grok 4.3's 1M) and cached input at $0.50, rising to $1.00 above 200K tokens. Provider-run benchmarks put it near Opus 4.8; independent neutral-harness runs put it well behind both Opus 4.8 and Fable 5 on coding, though token efficiency is truly strong.

DeepSeek V4 official launch confirmed for July 15, 2026 - The peak-hour pricing first announced June 30 is no longer theoretical - it goes live in two days. Off-peak V4 Flash and V4 Pro pricing is unchanged; peak-hour rates (9AM-noon and 2-6PM Beijing time) double as previously announced. See Hidden Costs for the full breakdown.

Hidden Costs

Grok 4.1 Fast Silently Repriced

Grok 4.1 Fast is the clearest hidden-cost trap in this table's history. xAI retired the grok-4-1-fast-reasoning and grok-4-1-fast-non-reasoning slugs on May 15, 2026, but instead of rejecting requests, both now silently redirect to Grok 4.3 - reasoning traffic maps to low reasoning effort, non-reasoning traffic maps to none. The API keeps responding normally, so nothing in application logs signals the switch. The billing consequence: input jumps from $0.20 to $1.25/MTok and output from $0.50 to $2.50/MTok, more than 6x on input. The slug retires completely on August 15, 2026, at which point requests will start failing outright. If any workload still targets the old slug, check the model field on your actual invoices - not the request payload - before assuming you're still paying legacy rates.

Claude Sonnet 5 Tokenizer Surcharge

Claude Sonnet 5 uses a newer tokenizer shared with Opus 4.7+, Fable 5, and Mythos 5. Anthropic's documentation notes roughly 30% more tokens for the same text compared to Sonnet 4.6 and earlier. At the $2/MTok intro rate, a workload costing $1.40 on Sonnet 4.6 ($3/MTok × typical 0.466 efficiency) costs ~$1.82 on Sonnet 5 after the tokenizer premium - still cheaper, but worth benchmarking with real prompt shapes.

After September 1, when the rate reverts to $3/MTok, the effective tokenizer-adjusted cost is ~$3.90/MTok equivalent - about 30% above Sonnet 4.6's standard rate. Teams planning long-term cost models on Sonnet 5 should include the tokenizer factor in projections.

Claude Opus 4.7 Fast Mode Deprecation

Anthropic is removing Fast Mode for Opus 4.7 on July 24, 2026. After that date, only Claude Opus 4.8 supports Fast Mode at $10/$50. Fast Mode for Opus 4.6 was already removed June 29, 2026. Applications currently routing Opus 4.7 with Fast Mode have roughly two weeks to migrate.

DeepSeek Peak-Hour Variable Rates

Effective July 15, 2026, V4 Flash peaks at $0.28/$0.56 and V4 Pro peaks at $0.87/$1.74 during Beijing business hours. The peak window is 9:00AM-noon and 2:00PM-6:00PM CST (UTC+8), which corresponds to 1-4AM and 6-10AM Eastern. Async batch workloads scheduled outside those windows are unaffected. DeepSeek says it will email affected accounts 24 hours before billing changes take effect. Synchronous production stacks in Asia-Pacific timezones should test actual billing before assuming off-peak rates.

GPT-5.5 Output Creep

GPT-5.5 input ($5/MTok) matches Claude Opus 4.8, but output ($30/MTok) runs 20% higher than Opus 4.8's $25/MTok. For reasoning-heavy or document-generation workloads producing long outputs, that difference compounds across millions of output tokens.

Rate Limits and Spend Tiers

OpenAI gates throughput by spend tier - new accounts cap at 500 RPM on frontier models; Tier 4 gets 10,000 RPM. Anthropic uses a four-tier structure. DeepSeek V4 Flash has no published tiers but queues under high load; latency spikes during the upcoming peak windows are likely to be more marked than during the current flat-rate period.

Batch API Discounts

OpenAI, Anthropic, Google, and xAI all offer 50% off async batch processing with 24-hour SLAs. Groq offers 50% off batch jobs. DeepSeek's automatic prompt caching at $0.0028 per cache-hit MTok competes with formal batch discounts without requiring a separate API endpoint. Claude Sonnet 5 batch pricing is $1/$5 during the intro period vs $1.50/$7.50 standard.

Prompt Caching

Cache hit pricing across major providers:

  • DeepSeek V4 Flash: $0.0028/MTok (98% off standard input)
  • DeepSeek V4 Pro: $0.003625/MTok (99.2% off list price)
  • Anthropic (all models): 10% of standard input ($0.20/MTok for Sonnet 5 intro, $0.50/MTok for Opus 4.8/4.7, $1.00/MTok for Fable 5)
  • GLM-5.2: $0.26/MTok (81% off standard input)
  • OpenAI: 10% of standard input (automatic, no setup)
  • Google: 10% of standard input plus storage fees ($0.15-$1.00/1M tokens/hour)
  • xAI Grok 4.5: $0.50/MTok cached (75% off), $1.00/MTok above 200K tokens
  • xAI Grok 4.3/4.20: $0.20/MTok cached (~84% off)
  • Groq: 50% off cached input tokens

Electronic shelf labels showing real-time price tags in a retail setting LLM API prices moved clearly in July 2026 - Sonnet 5 intro pricing undercuts Sonnet 4.6, Fable 5 returns from suspension, and DeepSeek announces peak-hour variable rates. Source: commons.wikimedia.org

Context Window Surcharges

Anthropic Fable 5, Opus 4.8/4.7, Sonnet 5, and Sonnet 4.6 include the full 1M context at standard rates with no surcharge. Gemini 3.1 Pro doubles input pricing above 200K tokens ($2 becomes $4, output goes $12 to $18). Gemini 2.5 Pro steps similarly ($1.25 to $2.50 above 200K). Grok 4.3 and Grok 4.20 offer 1M context at a flat $1.25/$2.50. GPT-5.6 Sol, Terra, and Luna all carry the same 1.05M context regardless of tier, with no surcharge tiering.

Grok 4.5 doubles its cached-input rate above 200K tokens ($0.50 to $1.00) despite a 500K context ceiling, and GPT-5.4 mini caps at 400K tokens. With Grok 4.1 Fast no longer pricing separately, GPT-5.6 Luna at $0.20/$1.25 with a full 1.05M context is now the most cost-efficient option in this table for long-context work without flagship pricing.

Free Tier Comparison

ProviderFree CreditsModels AvailableRate LimitsNotes
Google (Gemini)Unlimited free tierFlash-Lite, 2.5 Flash, 3.5 Flash5-15 RPM, 100-1,000 RPDPro models paid-only
GroqFree tier availableAll hosted modelsVaries by modelNo card required
xAI$175/month creditsAll Grok modelsStandard limitsVia data-sharing program
DeepSeek5M tokens on signupV4 Flash, V4 ProStandard limitsNon-renewable
OpenAI~$5 trial creditsGPT-4o mini, limited3 RPM (free tier)3-month expiry
Anthropic~$5 trial creditsAll modelsTier 1 limitsFew months expiry
MistralFree tier (limited)Ministral 3B legacyRate-limitedNo card required

Google's free tier remains the most useful for development - Flash models including Gemini 3.5 Flash are free with manageable rate limits. xAI's updated data-sharing program at $175/month in credits is now the most generous paid-alternative for API access without billing commitments. The new Claude Sonnet 5 intro pricing makes Anthropic's paid tier more competitive, but there's still no free development tier comparable to Google or Groq.

Price History

  • Jul 15, 2026 (scheduled) - DeepSeek V4 official release; peak-hour pricing (9AM-noon and 2-6PM CST) goes live. V4 Flash peaks at $0.28/$0.56, V4 Pro at $0.87/$1.74.

  • Jul 9, 2026 - GPT-5.6 Sol, Terra, and Luna reach general availability at $5/$30, $2.50/$15, and $1/$6. GPT-5.4 scheduled to retire July 23.

  • Jul 8, 2026 - Grok 4.5 launches at $2/$6 with a 500K context window, xAI's first full-scale V9-architecture model.

  • Jul 1, 2026 - Claude Fable 5 API access restored at $10/$50 after a 19-day suspension under US Commerce Department export controls. Subscription caps lifted July 7.

  • Jun 30, 2026 - Claude Sonnet 5 launches at $2/$10 introductory pricing through August 31, 2026 (then $3/$15). Uses newer tokenizer - ~30% more tokens vs Sonnet 4.6.

  • Jun 30, 2026 - DeepSeek announces peak-hour variable pricing for official mid-July V4 release. Rates double during 9AM-noon and 2-6PM CST.

  • Jun 16, 2026 - GLM-5.2 live on standalone Z.AI API at $1.40/$4.40 with 1M context. Cache hits at $0.26/MTok.

  • Jun 12, 2026 - Claude Fable 5 and Mythos 5 suspended under US export-control directive. API access fully halted for both models.

  • Jun 1, 2026 - DeepSeek V4 Pro promotional pricing ($0.435/$0.87) becomes permanent. The original list price of $1.74/$3.48 is retired.

  • May 28, 2026 - Claude Opus 4.8 launches at $5/$25 standard. Fast Mode drops from $30/$150 to $10/$50 vs Opus 4.7 - a 67% reduction.

  • May 15, 2026 - xAI retires eight legacy Grok model slugs, including Grok 4.1 Fast. Requests silently redirect to Grok 4.3 pricing ($1.25/$2.50) rather than erroring. Full slug shutdown Aug 15, 2026.

  • May 2026 - DeepSeek V4 Flash arrives on the API at $0.14/$0.28. Cache hits at $0.0028.

  • May 19, 2026 - Gemini 3.5 Flash launches at $1.50/$9.00. Batch at $0.75/$4.50.

  • May 7, 2026 - Gemini 3.1 Flash-Lite moves to GA at $0.25/$1.50.

  • Apr 2026 - Claude Opus 4.7 launches at $5/$25. Mistral overhauls the lineup: Mistral Small 4 at $0.10/$0.30, Mistral Small 3.2 at $0.08/$0.20.

Claude Sonnet 5's tokenizer produces ~30% more tokens vs Sonnet 4.6. At the $2/MTok intro rate you still come out ahead - but after September 1, the effective rate tops Sonnet 4.6 by about 30%.

FAQ

Which LLM API is cheapest per million tokens?

Ministral 3B at $0.04/$0.04 via the legacy endpoint is the cheapest standard commercial option. For production use, DeepSeek V4 Flash at $0.14/$0.28 off-peak is far more capable. DeepSeek's peak-hour rate doubling goes live July 15, 2026.

What's the best value LLM API for production right now?

DeepSeek V4 Flash at $0.14/$0.28 off-peak with 98% cache discounts at $0.0028/MTok. For frontier-class quality at a fair price, Claude Sonnet 5 at $2/$10 intro through August 31 is the best deal in the mid-tier.

Is GPT-5.6 available yet?

Yes. GPT-5.6 Sol, Terra, and Luna went generally available July 9, 2026, at $5/$30, $2.50/$15, and $1/$6 per MTok. GPT-5.4 retires July 23 and is replaced by Terra at the same price point.

Why did Grok 4.1 Fast disappear from this table?

It doesn't have its own pricing anymore. xAI retired the slug May 15, 2026, and requests now silently redirect to Grok 4.3 at $1.25/$2.50 - more than 6x the old input rate. The slug stops resolving completely August 15, 2026.

Will DeepSeek's peak-hour pricing affect my costs?

Only during Beijing business hours (9AM-noon and 2-6PM CST) once the official V4 launch takes effect July 15, 2026. For US-based workloads those windows correspond to 1-4AM and 6-10AM Eastern, easy to avoid with async batching. Real-time stacks in Asia-Pacific timezones may see costs double.

Are there free LLM APIs for development?

Google's Gemini API has the most useful free tier - Flash models including Gemini 3.5 Flash are free with rate limits. Groq provides free LPU inference on Llama and Qwen families. xAI offers $175/month in API credits via the data-sharing program. Mistral offers rate-limited access to the legacy Ministral 3B endpoint.


Sources:

✓ Last verified July 13, 2026

LinkedIn
Reddit
Hacker News
Telegram
LLM API Pricing Comparison - July 2026
About the author AI Benchmarks & Tools Analyst

James is a software engineer turned tech writer who spent six years building backend systems at a fintech startup in Chicago before pivoting to full-time analysis of AI tools and infrastructure.