LLM API Pricing Comparison - July 2026

Verified July 27: Claude Opus 5 lands at Opus 4.8 pricing ($5/$25) while claiming near-Fable-5 quality, Gemini 3.6 Flash cuts output pricing 17%, and Grok's whole lineup turns out to carry a 2x long-context surcharge nobody's headline price mentions.

Cheapest: Ministral 3B Best Value: DeepSeek V4 Flash Updated weekly
LLM API Pricing Comparison - July 2026

TL;DR

  • Claude Opus 5 launched July 24 at the same $5/$25 as Opus 4.8 - Anthropic says it closes in on Fable 5 quality at half the price
  • Gemini 3.6 Flash replaces 3.5 Flash as Google's workhorse tier - same $1.50 input, output cut 17% to $7.50
  • Every xAI Grok model carries a 2x surcharge on both input and output once a prompt crosses 200K tokens - a table showing only the base rate understates real cost on long-context jobs
  • GPT-5.4's reported July 23 retirement never happened - it's still live on OpenAI's pricing page with no deprecation date

The Bottom Line

Claude Opus 5 is the story this week. Anthropic shipped it July 24 at unchanged pricing from Opus 4.8 - still $5/$25 per million tokens, still $10/$50 under Fast Mode - while claiming it closes most of the gap to Fable 5 on coding and agentic benchmarks. If that quality claim holds up under independent testing, Opus 5 becomes the best value pick at the frontier tier: same price as its predecessor, no new premium for the upgrade. Opus 4.8 stays available alongside it rather than being retired outright.

The same day, Anthropic finished a change flagged as pending in this table two weeks ago: Claude Opus 4.7 lost Fast Mode access completely on July 24. Only Opus 5 and Opus 4.8 support it now, at $10/$50.

Google shipped Gemini 3.6 Flash on July 21, with a new Gemini 3.5 Flash-Lite tier. 3.6 Flash keeps 3.5 Flash's $1.50 input price but cuts output from $9.00 to $7.50 - a 17% reduction - while independent benchmarking firm Artificial Analysis measured a 18% drop in average task cost thanks to shorter completions. 3.5 Flash-Lite slots in at $0.30/$2.50, replacing Gemini 2.5 Flash-Lite as Google's cheapest paid tier.

The correction worth flagging: last update reported GPT-5.4 retiring July 23 with traffic rolling onto GPT-5.6 Terra. That retirement didn't happen. OpenAI's official pricing and deprecation pages both still list GPT-5.4, GPT-5.4-mini, and GPT-5.4-nano as active with no shutdown date. Terra and GPT-5.4 now simply coexist at the same $2.50/$15 price point - if you built a migration around the July 23 date, there's nothing to migrate yet.

The other correction is on xAI pricing, and it's been wrong in this table for longer. Every current Grok model - Grok 4.5, Grok 4.3, Grok 4.20, and Grok Build 0.1 - bills at a flat rate only below 200K prompt tokens. Cross that threshold and both input and output double for the entire request, not just the overage. xAI's own docs state it plainly: once a prompt reaches the threshold, the higher rate applies to all tokens in that request. The table below has always shown the under-200K rate; the row notes now carry the long-context number too.

DeepSeek's peak-hour pricing is no longer a scheduled change - it's been live since July 15. Off-peak V4 Flash stays at $0.14/$0.28; peak-hour (9AM-noon and 2-6PM Beijing time) runs $0.28/$0.56. DeepSeek V4 Flash remains the best raw value for workloads that can run off-peak or stay in US timezones.

Full Pricing Table

All prices in USD per million tokens (MTok). Verified against official documentation July 27, 2026. Sorted by input price, cheapest first.

ModelProviderInput (/1M)Output (/1M)ContextNotes
Ministral 3BMistral$0.04$0.04256KLegacy endpoint only
Llama 3.1 8BGroq$0.05$0.08128K840+ tok/s on LPU
GPT OSS 20BGroq$0.075$0.30128KOpen-weight on LPU
Mistral Small 3.2Mistral$0.08$0.20128KCheapest Mistral with function calling
GPT-4.1 nanoOpenAI$0.10$0.401MRouting and classification
Gemini 2.5 Flash-LiteGoogle$0.10$0.401MFree tier
Mistral Small 4Mistral$0.10$0.30128K
Ministral 3 3B 2512Mistral$0.10$0.10256KCurrent gen; legacy 3B endpoint at $0.04
Llama 4 ScoutGroq$0.11$0.34128K17B active / 16E MoE
DeepSeek V4 FlashDeepSeek$0.14$0.281MCache hit $0.0028; best value - peak-hour doubling live since Jul 15
Ministral 3-8BMistral$0.15$0.15256K
GPT OSS 120BGroq$0.15$0.60128K
GPT-5.6 LunaOpenAI$0.20$1.251.05MCached $0.02
GPT-5.4 nanoOpenAI$0.20$1.25n/aCached input $0.02; still active, no retirement date
Ministral 3-14BMistral$0.20$0.20256KFlat in/out; optional reasoning mode
Gemini 3.1 Flash-LiteGoogle$0.25$1.501MFree tier
Gemini 3.5 Flash-LiteGoogle$0.30$2.501MNew Jul 21; free tier; replaces 2.5 Flash-Lite as Google's cheapest paid tier
Qwen3 32BGroq$0.29$0.59131KOpen-weight; strong multilingual
Gemini 2.5 FlashGoogle$0.30$2.501MFree tier; solid mid-range
GPT-4.1 miniOpenAI$0.40$1.601M
Devstral 2Mistral$0.40$2.00256KCoding/dev agent specialized
DeepSeek V4 ProDeepSeek$0.435$0.871MCache hit $0.003625; peak-hour doubling live since Jul 15
Mistral Large 3Mistral$0.50$1.50262K75% below Mistral Large 2
Qwen 3.6 27BGroq$0.60$3.00131KNewest Groq-hosted open-weight addition
Llama 3.3 70BGroq$0.59$0.79128KDense 70B; instruction following
GPT-5.4 miniOpenAI$0.75$4.50400KCached $0.075; still active, no retirement date despite prior reports
Claude Haiku 4.5Anthropic$1.00$5.00200KCheapest active Anthropic model
Grok Build 0.1xAI$1.00$2.00256K$2.00/$4.00 at or above 200K prompt tokens
o4-miniOpenAI$1.10$4.40200KCheapest dedicated reasoning model
Grok 4.3xAI$1.25$2.501M$2.50/$5.00 at or above 200K prompt tokens
Grok 4.20xAI$1.25$2.501MReasoning variant; same tiered pricing as 4.3
Gemini 2.5 ProGoogle$1.25$10.001M$2.50/$15 above 200K tokens
GPT-5.1OpenAI$1.25$10.00128KShuts down Aug 10, 2026; migrate to GPT-5.6 Sol
GLM-5.2Z.AI$1.40$4.401MOpen weights; cache hit $0.26
Gemini 3.6 FlashGoogle$1.50$7.501MNew Jul 21; replaces 3.5 Flash as Google's workhorse tier
Gemini 3.5 FlashGoogle$1.50$9.001MFree tier; batch $0.75/$4.50; still available with 3.6 Flash
Mistral Medium 3.5Mistral$1.50$7.50128K
GPT-5.2OpenAI$1.75$14.00128KShuts down Aug 10, 2026; migrate to GPT-5.6 Sol
GPT-5.3-codexOpenAI$1.75$14.00128KCode specialist
o3OpenAI$2.00$8.00200K80% cut from original o1 pricing
GPT-4.1OpenAI$2.00$8.001M
Gemini 3.1 Pro PreviewGoogle$2.00$12.001M$4/$18 above 200K tokens
Claude Sonnet 5Anthropic$2.00$10.001MIntro through Aug 31; $3/$15 after; new tokenizer adds ~30% tokens
Grok 4.5xAI$2.00$6.00500K$4.00/$12.00 at or above 200K prompt tokens; cached $0.50
GPT-5.6 TerraOpenAI$2.50$15.001.05MSame price as GPT-5.4, which didn't retire as previously reported
GPT-5.4OpenAI$2.50$15.001MStill active - the reported Jul 23 retirement did not happen
Claude Sonnet 4.6Anthropic$3.00$15.001M
Claude Opus 4.7Anthropic$5.00$25.001MFast Mode removed Jul 24
Claude Opus 4.8Anthropic$5.00$25.001MFast Mode $10/$50; stays available with Opus 5
Claude Opus 5Anthropic$5.00$25.001MNew Jul 24; same price as Opus 4.8, claims near-Fable-5 quality; Fast Mode $10/$50
GPT-5.6 SolOpenAI$5.00$30.001.05MMatches GPT-5.5 price
GPT-5.5OpenAI$5.00$30.001MStays available alongside GPT-5.6
Claude Fable 5Anthropic$10.00$50.001MRestored Jul 1; suspended Jun 12-Jun 30
GPT-5.4 Pro / GPT-5.5 ProOpenAI$30.00$180.001MResearch/enterprise tier

Grok 4.1 Fast is omitted above because it no longer has its own price - see "Grok 4.1 Fast Silently Repriced" under Hidden Costs.

For benchmark context behind these prices, see the cost-efficiency leaderboard.

A calculator and pen on a notebook, representing the cost math behind API pricing decisions Working out real API costs requires more than reading the headline price - caching, batching, context surcharges, and tokenizer differences all shift the final number. Source: unsplash.com

Changes Since July 13

Five standout changes since the last update.

Claude Opus 5 launches (July 24, 2026) - Claude Opus 5 shipped at unchanged pricing from Opus 4.8 - $5/$25 standard, $10/$50 under Fast Mode. Anthropic's own benchmarks put it near Fable 5 on coding and agentic tasks; independent verification of that claim is still pending. Opus 4.8 remains available rather than being retired, so buyers effectively get to choose between the two at identical prices until real-world benchmarks settle which is actually better for a given workload.

GPT-5.4's reported retirement didn't happen - The last update in this table said GPT-5.4 would retire July 23 with traffic rolling onto GPT-5.6 Terra. Checking OpenAI's official pricing and deprecation pages now, as of July 27 GPT-5.4, GPT-5.4-mini, and GPT-5.4-nano are all still listed as active with no shutdown date. Terra and GPT-5.4 simply coexist at the same $2.50/$15 price. This is a correction to prior coverage, not a new announcement - treat any "GPT-5.4 is gone" claim you see elsewhere as outdated.

xAI's Grok lineup carries a long-context surcharge this table previously missed - Grok 4.5, Grok 4.3, Grok 4.20, and Grok Build 0.1 all bill at 2x both input and output once a prompt reaches 200K tokens, applied to the whole request rather than just the tokens over the line. xAI's docs are explicit about this. Prior editions of this table only showed the sub-200K rate; see Hidden Costs below for the corrected numbers.

Gemini 3.6 Flash and 3.5 Flash-Lite launch (July 21, 2026) - Gemini 3.6 Flash keeps 3.5 Flash's $1.50 input but cuts output 17% to $7.50, and Google says it also produces 17% fewer output tokens per task, compounding the savings. Gemini 3.5 Flash-Lite replaces 2.5 Flash-Lite as the cheapest paid Gemini tier at $0.30/$2.50.

DeepSeek V4 peak-hour pricing is now confirmed live - The doubling first announced June 30 and scheduled for July 15 has been in effect for nearly two weeks. Off-peak V4 Flash and V4 Pro pricing is unchanged; peak-hour rates (9AM-noon and 2-6PM Beijing time) run double the off-peak rate. See Hidden Costs for the full breakdown.

Hidden Costs

Grok's Long-Context Surcharge Doubles the Whole Bill

Every current Grok model bills a flat rate below 200K prompt tokens and exactly 2x that rate - on both input and output, for the entire request - once the prompt crosses the threshold. Grok 4.5 goes from $2/$6 to $4/$12, Grok 4.3 and Grok 4.20 go from $1.25/$2.50 to $2.50/$5.00, and Grok Build 0.1 goes from $1/$2 to $2/$4. This isn't a marginal surcharge on the overage - xAI's documentation states the higher rate applies to all tokens in a request that crosses the line, so a 199K-token prompt and a 201K-token prompt can bill more than 2x apart in total cost. Workloads that occasionally spike past 200K tokens (long documents, large codebases, multi-turn agent sessions) should budget for the high-tier rate rather than assume the headline price.

Grok 4.1 Fast's Final Weeks

Grok 4.1 Fast hasn't billed at its old $0.20/$0.50 rate since May 15, 2026 - xAI retired the slug and silently redirects requests to Grok 4.3 pricing instead of erroring. The slug stops resolving completely on August 15, 2026, about three weeks from this update. Anyone still targeting the old model ID should migrate now rather than wait for it to start failing.

Claude Sonnet 5 Tokenizer Surcharge

Claude Sonnet 5 uses a newer tokenizer shared with Opus 4.7+, Fable 5, and Mythos 5. Anthropic's documentation notes roughly 30% more tokens for the same text compared to Sonnet 4.6 and earlier. At the $2/MTok intro rate, a workload costing $1.40 on Sonnet 4.6 ($3/MTok × typical 0.466 efficiency) costs ~$1.82 on Sonnet 5 after the tokenizer premium - still cheaper, but worth benchmarking with real prompt shapes.

After September 1, when the rate reverts to $3/MTok, the effective tokenizer-adjusted cost is ~$3.90/MTok equivalent - about 30% above Sonnet 4.6's standard rate. Teams planning long-term cost models on Sonnet 5 should include the tokenizer factor in projections.

Claude Opus 4.7 Lost Fast Mode

Anthropic removed Fast Mode for Opus 4.7 on July 24, 2026, the same day Opus 5 launched. Only Claude Opus 5 and Claude Opus 4.8 support Fast Mode now, both at $10/$50. Fast Mode for Opus 4.6 was already removed June 29, 2026. Any application still requesting speed: "fast" on Opus 4.7 now gets an error instead of a response.

DeepSeek Peak-Hour Variable Rates

Live since July 15, 2026: V4 Flash peaks at $0.28/$0.56 and V4 Pro peaks at $0.87/$1.74 during Beijing business hours. The peak window is 9:00AM-noon and 2:00PM-6:00PM CST (UTC+8), which corresponds to 1-4AM and 6-10AM Eastern. Async batch workloads scheduled outside those windows are unaffected. Synchronous production stacks in Asia-Pacific timezones should verify actual billing rather than assume off-peak rates.

GPT-5.5 Output Creep

GPT-5.5 input ($5/MTok) matches Claude Opus 4.8 and Claude Opus 5, but output ($30/MTok) runs 20% higher than either Opus model's $25/MTok. For reasoning-heavy or document-generation workloads producing long outputs, that difference compounds across millions of output tokens.

Rate Limits and Spend Tiers

OpenAI gates throughput by spend tier - new accounts cap at 500 RPM on frontier models; Tier 4 gets 10,000 RPM. Anthropic uses a four-tier structure. DeepSeek V4 Flash has no published tiers but queues under high load; latency spikes during the upcoming peak windows are likely to be more marked than during the current flat-rate period.

Batch API Discounts

OpenAI, Anthropic, Google, and xAI all offer 50% off async batch processing with 24-hour SLAs. Groq offers 50% off batch jobs. DeepSeek's automatic prompt caching at $0.0028 per cache-hit MTok competes with formal batch discounts without requiring a separate API endpoint. Claude Sonnet 5 batch pricing is $1/$5 during the intro period vs $1.50/$7.50 standard.

Prompt Caching

Cache hit pricing across major providers:

  • DeepSeek V4 Flash: $0.0028/MTok (98% off standard input)
  • DeepSeek V4 Pro: $0.003625/MTok (99.2% off list price)
  • Anthropic (all models): 10% of standard input ($0.20/MTok for Sonnet 5 intro, $0.50/MTok for Opus 4.8/4.7, $1.00/MTok for Fable 5)
  • GLM-5.2: $0.26/MTok (81% off standard input)
  • OpenAI: 10% of standard input (automatic, no setup)
  • Google: 10% of standard input plus storage fees ($0.15-$1.00/1M tokens/hour)
  • xAI Grok 4.5: $0.50/MTok cached (75% off) below 200K tokens; base input/output rates double above the threshold - see Hidden Costs above
  • xAI Grok 4.3/4.20: $0.20/MTok cached (~84% off) below 200K tokens
  • Groq: 50% off cached input tokens

Electronic shelf labels showing real-time price tags in a retail setting LLM API prices moved clearly in July 2026 - Sonnet 5 intro pricing undercuts Sonnet 4.6, Fable 5 returns from suspension, and DeepSeek announces peak-hour variable rates. Source: commons.wikimedia.org

Context Window Surcharges

Anthropic Fable 5, Opus 5, Opus 4.8/4.7, Sonnet 5, and Sonnet 4.6 include the full 1M context at standard rates with no surcharge. Gemini 3.1 Pro doubles input pricing above 200K tokens ($2 becomes $4, output goes $12 to $18). Gemini 2.5 Pro steps similarly ($1.25 to $2.50 above 200K). GPT-5.6 Sol, Terra, and Luna all carry the same 1.05M context regardless of tier, with no surcharge tiering.

Every current Grok model doubles its rate above 200K prompt tokens - not just the cached-input rate, the full input and output price. Grok 4.5 goes to $4/$12 despite a 500K context ceiling, and Grok 4.3/Grok 4.20 go to $2.50/$5.00 within their 1M window. GPT-5.4 mini caps at 400K tokens with no surcharge tier of its own. With Grok 4.1 Fast no longer pricing separately, GPT-5.6 Luna at $0.20/$1.25 with a full 1.05M context and no surcharge tiering is the most cost-efficient option in this table for long-context work without flagship pricing.

Free Tier Comparison

ProviderFree CreditsModels AvailableRate LimitsNotes
Google (Gemini)Unlimited free tierFlash-Lite, 2.5 Flash, 3.5 Flash, 3.6 Flash5-15 RPM, 100-1,000 RPDPro models paid-only
GroqFree tier availableAll hosted modelsVaries by modelNo card required
xAI$175/month creditsAll Grok modelsStandard limitsVia data-sharing program
DeepSeek5M tokens on signupV4 Flash, V4 ProStandard limitsNon-renewable
OpenAI~$5 trial creditsGPT-4o mini, limited3 RPM (free tier)3-month expiry
Anthropic~$5 trial creditsAll modelsTier 1 limitsFew months expiry
MistralFree tier (limited)Ministral 3B legacyRate-limitedNo card required

Google's free tier remains the most useful for development - Flash models including Gemini 3.6 Flash are free with manageable rate limits. xAI's updated data-sharing program at $175/month in credits is now the most generous paid-alternative for API access without billing commitments. The new Claude Sonnet 5 intro pricing makes Anthropic's paid tier more competitive, but there's still no free development tier comparable to Google or Groq.

Price History

  • Jul 24, 2026 - Claude Opus 5 launches at unchanged $5/$25 pricing from Opus 4.8; same day, Claude Opus 4.7 loses Fast Mode access.

  • Jul 21, 2026 - Gemini 3.6 Flash launches at $1.50/$7.50 (output down 17% from 3.5 Flash); Gemini 3.5 Flash-Lite launches at $0.30/$2.50.

  • Jul 15, 2026 - DeepSeek V4 official release; peak-hour pricing (9AM-noon and 2-6PM CST) goes live. V4 Flash peaks at $0.28/$0.56, V4 Pro at $0.87/$1.74.

  • Jul 9, 2026 - GPT-5.6 Sol, Terra, and Luna reach general availability at $5/$30, $2.50/$15, and $1/$6. GPT-5.4's July 23 retirement was reported but didn't happen - it remains active.

  • Jul 8, 2026 - Grok 4.5 launches at $2/$6 with a 500K context window, xAI's first full-scale V9-architecture model.

  • Jul 1, 2026 - Claude Fable 5 API access restored at $10/$50 after a 19-day suspension under US Commerce Department export controls. Subscription caps lifted July 7.

  • Jun 30, 2026 - Claude Sonnet 5 launches at $2/$10 introductory pricing through August 31, 2026 (then $3/$15). Uses newer tokenizer - ~30% more tokens vs Sonnet 4.6.

  • Jun 30, 2026 - DeepSeek announces peak-hour variable pricing for official mid-July V4 release. Rates double during 9AM-noon and 2-6PM CST.

  • Jun 16, 2026 - GLM-5.2 live on standalone Z.AI API at $1.40/$4.40 with 1M context. Cache hits at $0.26/MTok.

  • Jun 12, 2026 - Claude Fable 5 and Mythos 5 suspended under US export-control directive. API access fully halted for both models.

  • Jun 1, 2026 - DeepSeek V4 Pro promotional pricing ($0.435/$0.87) becomes permanent. The original list price of $1.74/$3.48 is retired.

  • May 28, 2026 - Claude Opus 4.8 launches at $5/$25 standard. Fast Mode drops from $30/$150 to $10/$50 vs Opus 4.7 - a 67% reduction.

  • May 15, 2026 - xAI retires eight legacy Grok model slugs, including Grok 4.1 Fast. Requests silently redirect to Grok 4.3 pricing ($1.25/$2.50) rather than erroring. Full slug shutdown Aug 15, 2026.

  • May 2026 - DeepSeek V4 Flash arrives on the API at $0.14/$0.28. Cache hits at $0.0028.

  • May 19, 2026 - Gemini 3.5 Flash launches at $1.50/$9.00. Batch at $0.75/$4.50.

  • May 7, 2026 - Gemini 3.1 Flash-Lite moves to GA at $0.25/$1.50.

  • Apr 2026 - Claude Opus 4.7 launches at $5/$25. Mistral overhauls the lineup: Mistral Small 4 at $0.10/$0.30, Mistral Small 3.2 at $0.08/$0.20.

Claude Sonnet 5's tokenizer produces ~30% more tokens vs Sonnet 4.6. At the $2/MTok intro rate you still come out ahead - but after September 1, the effective rate tops Sonnet 4.6 by about 30%.

FAQ

Which LLM API is cheapest per million tokens?

Ministral 3B at $0.04/$0.04 via the legacy endpoint is the cheapest standard commercial option. For production use, DeepSeek V4 Flash at $0.14/$0.28 off-peak is far more capable. Peak-hour rate doubling has been live since July 15, 2026.

What's the best value LLM API for production right now?

DeepSeek V4 Flash at $0.14/$0.28 off-peak with 98% cache discounts at $0.0028/MTok. For frontier-class quality, Claude Opus 5 at $5/$25 claims near-Fable-5 results at half the Fable 5 price - the best value pick at the top tier if that benchmark claim holds up independently.

Is GPT-5.4 still available?

Yes. Despite earlier reports of a July 23, 2026 retirement, GPT-5.4, GPT-5.4-mini, and GPT-5.4-nano remain listed on OpenAI's official pricing page with no deprecation date as of July 27, 2026. It now coexists with GPT-5.6 Terra at the same $2.50/$15 price.

Does Grok pricing really double on long prompts?

Yes. Every current Grok model - 4.5, 4.3, 4.20, and Build 0.1 - bills at 2x the base input and output rate once a prompt reaches 200K tokens, applied to the entire request. Grok 4.5 goes from $2/$6 to $4/$12 above that threshold.

Why did Grok 4.1 Fast disappear from this table?

It doesn't have its own pricing anymore. xAI retired the slug May 15, 2026, and requests now silently redirect to Grok 4.3 at $1.25/$2.50 - more than 6x the old input rate. The slug stops resolving completely August 15, 2026.

Are there free LLM APIs for development?

Google's Gemini API has the most useful free tier - Flash models including Gemini 3.6 Flash are free with rate limits. Groq provides free LPU inference on Llama and Qwen families. xAI offers $175/month in API credits via the data-sharing program. Mistral offers rate-limited access to the legacy Ministral 3B endpoint.


Sources:

✓ Last verified July 27, 2026

LinkedIn
Reddit
Hacker News
Telegram
LLM API Pricing Comparison - July 2026
About the author AI Benchmarks & Tools Analyst

James is a software engineer turned tech writer who spent six years building backend systems at a fintech startup in Chicago before pivoting to full-time analysis of AI tools and infrastructure.