LLM API Pricing Comparison - July 2026
Verified July 13: GPT-5.6 Sol/Terra/Luna hit GA at $5/$30, $2.50/$15 and $1/$6, Grok 4.5 lands at $2/$6, Grok 4.1 Fast quietly redirects to pricier Grok 4.3, DeepSeek's peak-hour doubling goes live July 15.

TL;DR
- GPT-5.6 Sol/Terra/Luna went GA July 9 - Sol matches GPT-5.5 at $5/$30, Terra matches the soon-retired GPT-5.4 at $2.50/$15, Luna opens a new $1/$6 tier
- Grok 4.1 Fast is no longer the cheap long-context pick - the slug quietly redirects to Grok 4.3 pricing ($1.25/$2.50) since May 15 and fully retires August 15
- Grok 4.5 launched at $2/$6 with a 500K context window, undercutting GPT-5.4 mini on output while beating it on quality claims
The Bottom Line
GPT-5.6 is the headline change this week. OpenAI previewed the three-tier family - Sol, Terra, Luna - to a small government-vetted group on June 26, then took it fully generally available on July 9 after clearing an additional round of review tied to the administration's AI cybersecurity order. Pricing lands exactly where the preview numbers pointed: Sol at $5/$30 (matching GPT-5.5), Terra at $2.50/$15 (matching GPT-5.4), and Luna as a new bargain tier at $1/$6. All three carry a 1.05M token context window, 128K max output, and cached input at 10% of the standard rate. GPT-5.4 itself retires July 23, so Terra is effectively its replacement at the same price.
The more consequential correction this week is on the cheap end of the xAI lineup. Grok 4.1 Fast has been listed in this table for months at $0.20/$0.50 with a 2M context window - and that pricing hasn't been real since May 15, 2026. xAI retired the slug that day along with seven other legacy models; requests still resolve without erroring, but they silently redirect to Grok 4.3 and bill at $1.25/$2.50, more than 6x the old input rate. The slug disappears completely on August 15. Anyone still routing to grok-4-1-fast-reasoning or -non-reasoning in production has been overpaying for two months without a code change to show for it - check your xAI invoices against the model field, not just the slug you sent.
Grok 4.5 is the actual new xAI model worth assessing. It launched publicly July 8 at $2/$6 with a 500K context window (down from Grok 4.3's 1M) and cached input at $0.50 - or $1.00 above 200K tokens. It's a 1.5T-parameter MoE model, xAI's first V9-architecture model at full scale, and token-efficient in practice: independent benchmarks put it well behind Opus 4.8 and Fable 5 on coding quality despite provider-run harnesses showing closer results.
DeepSeek's peak-hour pricing is no longer a future announcement - the official V4 release lands July 15, two days after this update. Off-peak V4 Flash stays at $0.14/$0.28; peak-hour (9AM-noon and 2-6PM Beijing time) rises to $0.28/$0.56. DeepSeek V4 Flash remains the best raw value for workloads that can run off-peak or stay in US timezones.
Full Pricing Table
All prices in USD per million tokens (MTok). Verified against official documentation July 13, 2026. Sorted by input price, cheapest first.
| Model | Provider | Input (/1M) | Output (/1M) | Context | Notes |
|---|---|---|---|---|---|
| Ministral 3B | Mistral | $0.04 | $0.04 | 256K | Legacy endpoint only |
| Llama 3.1 8B | Groq | $0.05 | $0.08 | 128K | 840+ tok/s on LPU |
| GPT OSS 20B | Groq | $0.075 | $0.30 | 128K | Open-weight on LPU |
| Mistral Small 3.2 | Mistral | $0.08 | $0.20 | 128K | Cheapest Mistral with function calling |
| GPT-4.1 nano | OpenAI | $0.10 | $0.40 | 1M | Routing and classification |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | Free tier | |
| Mistral Small 4 | Mistral | $0.10 | $0.30 | 128K | |
| Ministral 3 3B 2512 | Mistral | $0.10 | $0.10 | 256K | Current gen; legacy 3B endpoint at $0.04 |
| Llama 4 Scout | Groq | $0.11 | $0.34 | 128K | 17B active / 16E MoE |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | 1M | Cache hit $0.0028; best value - peak-hour doubling goes official Jul 15 |
| Ministral 3-8B | Mistral | $0.15 | $0.15 | 256K | |
| GPT OSS 120B | Groq | $0.15 | $0.60 | 128K | |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.25 | 1.05M | New Jul 9 GA; cached $0.02 |
| GPT-5.4 nano | OpenAI | $0.20 | $1.25 | n/a | Cached input $0.02 |
| Ministral 3-14B | Mistral | $0.20 | $0.20 | 256K | Flat in/out; optional reasoning mode |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M | Free tier | |
| Qwen3 32B | Groq | $0.29 | $0.59 | 131K | Open-weight; strong multilingual |
| Gemini 2.5 Flash | $0.30 | $2.50 | 1M | Free tier; solid mid-range | |
| GPT-4.1 mini | OpenAI | $0.40 | $1.60 | 1M | |
| Devstral 2 | Mistral | $0.40 | $2.00 | 256K | Coding/dev agent specialized |
| DeepSeek V4 Pro | DeepSeek | $0.435 | $0.87 | 1M | Cache hit $0.003625; peak-hour doubling goes official Jul 15 |
| Mistral Large 3 | Mistral | $0.50 | $1.50 | 262K | 75% below Mistral Large 2 |
| Llama 3.3 70B | Groq | $0.59 | $0.79 | 128K | Dense 70B; instruction following |
| GPT-5.4 mini | OpenAI | $0.75 | $4.50 | 400K | Cached $0.075; retires with GPT-5.4 Jul 23 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | 200K | Cheapest active Anthropic model |
| Grok Build 0.1 | xAI | $1.00 | $2.00 | 256K | Agentic coding; native MCP |
| o4-mini | OpenAI | $1.10 | $4.40 | 200K | Cheapest dedicated reasoning model |
| Grok 4.3 | xAI | $1.25 | $2.50 | 1M | Now also the effective price for redirected Grok 4.1 Fast traffic |
| Grok 4.20 | xAI | $1.25 | $2.50 | 1M | Reasoning variant; same price as 4.3 |
| Gemini 2.5 Pro | $1.25 | $10.00 | 1M | $2.50/$15 above 200K tokens | |
| GPT-5.1 | OpenAI | $1.25 | $10.00 | 128K | |
| GLM-5.2 | Z.AI | $1.40 | $4.40 | 1M | Open weights; cache hit $0.26 |
| Gemini 3.5 Flash | $1.50 | $9.00 | 1M | Free tier; batch $0.75/$4.50 | |
| Mistral Medium 3.5 | Mistral | $1.50 | $7.50 | 128K | |
| GPT-5.2 | OpenAI | $1.75 | $14.00 | 128K | |
| GPT-5.3-codex | OpenAI | $1.75 | $14.00 | 128K | Code specialist |
| o3 | OpenAI | $2.00 | $8.00 | 200K | 80% cut from original o1 pricing |
| GPT-4.1 | OpenAI | $2.00 | $8.00 | 1M | |
| Gemini 3.1 Pro Preview | $2.00 | $12.00 | 1M | $4/$18 above 200K tokens | |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | 1M | Intro through Aug 31; $3/$15 after; new tokenizer adds ~30% tokens |
| Grok 4.5 | xAI | $2.00 | $6.00 | 500K | New Jul 8; cached $0.50 ($1.00 above 200K) |
| GPT-5.6 Terra | OpenAI | $2.50 | $15.00 | 1.05M | New Jul 9 GA; matches retiring GPT-5.4 price |
| GPT-5.4 | OpenAI | $2.50 | $15.00 | 1M | Retires Jul 23; replaced by GPT-5.6 Terra |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | 1M | |
| Claude Opus 4.7 | Anthropic | $5.00 | $25.00 | 1M | Fast Mode deprecated Jul 24 |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | 1M | Fast Mode $10/$50; standard Anthropic flagship |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | 1.05M | New Jul 9 GA; matches GPT-5.5 price |
| GPT-5.5 | OpenAI | $5.00 | $30.00 | 1M | Stays available alongside GPT-5.6 |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | 1M | Restored Jul 1; suspended Jun 12-Jun 30 |
| GPT-5.4 Pro / GPT-5.5 Pro | OpenAI | $30.00 | $180.00 | 1M | Research/enterprise tier |
Grok 4.1 Fast is omitted above because it no longer has its own price - see "Grok 4.1 Fast Silently Repriced" under Hidden Costs.
For benchmark context behind these prices, see the cost-efficiency leaderboard.
Working out real API costs requires more than reading the headline price - caching, batching, context surcharges, and tokenizer differences all shift the final number.
Source: unsplash.com
Changes Since July 6
Four standout changes this week.
GPT-5.6 goes GA (July 9, 2026) - GPT-5.6 Sol, Terra, and Luna cleared the small-partner preview and shipped to everyone July 9, three days after this table was last verified. Pricing matches what leaked during preview: Sol $5/$30 (same as GPT-5.5), Terra $2.50/$15 (same as the outgoing GPT-5.4), Luna $1/$6 as a new budget tier with a full 1.05M context window. Cached input is 10% of standard across all three. GPT-5.4 itself retires July 23 and its traffic rolls onto Terra at the same price - a rare same-cost forced migration.
Grok 4.1 Fast quietly repriced (effective since May 15, 2026) - This is a correction, not new news: Grok 4.1 Fast has not actually billed at its listed $0.20/$0.50 rate since May 15. xAI retired the slug that day along with seven other legacy models and pointed all requests to Grok 4.3 pricing ($1.25/$2.50) instead of returning an error. We're pulling the stale row from this table now and flagging it under Hidden Costs below - anyone who copied last month's numbers into a cost model has been under-forecasting by more than 6x on input.
Grok 4.5 launches (July 8, 2026) - Grok 4.5 is xAI's first V9-architecture model at full scale, a 1.5T-parameter MoE priced at $2/$6 with a 500K context window (down from Grok 4.3's 1M) and cached input at $0.50, rising to $1.00 above 200K tokens. Provider-run benchmarks put it near Opus 4.8; independent neutral-harness runs put it well behind both Opus 4.8 and Fable 5 on coding, though token efficiency is truly strong.
DeepSeek V4 official launch confirmed for July 15, 2026 - The peak-hour pricing first announced June 30 is no longer theoretical - it goes live in two days. Off-peak V4 Flash and V4 Pro pricing is unchanged; peak-hour rates (9AM-noon and 2-6PM Beijing time) double as previously announced. See Hidden Costs for the full breakdown.
Hidden Costs
Grok 4.1 Fast Silently Repriced
Grok 4.1 Fast is the clearest hidden-cost trap in this table's history. xAI retired the grok-4-1-fast-reasoning and grok-4-1-fast-non-reasoning slugs on May 15, 2026, but instead of rejecting requests, both now silently redirect to Grok 4.3 - reasoning traffic maps to low reasoning effort, non-reasoning traffic maps to none. The API keeps responding normally, so nothing in application logs signals the switch. The billing consequence: input jumps from $0.20 to $1.25/MTok and output from $0.50 to $2.50/MTok, more than 6x on input. The slug retires completely on August 15, 2026, at which point requests will start failing outright. If any workload still targets the old slug, check the model field on your actual invoices - not the request payload - before assuming you're still paying legacy rates.
Claude Sonnet 5 Tokenizer Surcharge
Claude Sonnet 5 uses a newer tokenizer shared with Opus 4.7+, Fable 5, and Mythos 5. Anthropic's documentation notes roughly 30% more tokens for the same text compared to Sonnet 4.6 and earlier. At the $2/MTok intro rate, a workload costing $1.40 on Sonnet 4.6 ($3/MTok × typical 0.466 efficiency) costs ~$1.82 on Sonnet 5 after the tokenizer premium - still cheaper, but worth benchmarking with real prompt shapes.
After September 1, when the rate reverts to $3/MTok, the effective tokenizer-adjusted cost is ~$3.90/MTok equivalent - about 30% above Sonnet 4.6's standard rate. Teams planning long-term cost models on Sonnet 5 should include the tokenizer factor in projections.
Claude Opus 4.7 Fast Mode Deprecation
Anthropic is removing Fast Mode for Opus 4.7 on July 24, 2026. After that date, only Claude Opus 4.8 supports Fast Mode at $10/$50. Fast Mode for Opus 4.6 was already removed June 29, 2026. Applications currently routing Opus 4.7 with Fast Mode have roughly two weeks to migrate.
DeepSeek Peak-Hour Variable Rates
Effective July 15, 2026, V4 Flash peaks at $0.28/$0.56 and V4 Pro peaks at $0.87/$1.74 during Beijing business hours. The peak window is 9:00AM-noon and 2:00PM-6:00PM CST (UTC+8), which corresponds to 1-4AM and 6-10AM Eastern. Async batch workloads scheduled outside those windows are unaffected. DeepSeek says it will email affected accounts 24 hours before billing changes take effect. Synchronous production stacks in Asia-Pacific timezones should test actual billing before assuming off-peak rates.
GPT-5.5 Output Creep
GPT-5.5 input ($5/MTok) matches Claude Opus 4.8, but output ($30/MTok) runs 20% higher than Opus 4.8's $25/MTok. For reasoning-heavy or document-generation workloads producing long outputs, that difference compounds across millions of output tokens.
Rate Limits and Spend Tiers
OpenAI gates throughput by spend tier - new accounts cap at 500 RPM on frontier models; Tier 4 gets 10,000 RPM. Anthropic uses a four-tier structure. DeepSeek V4 Flash has no published tiers but queues under high load; latency spikes during the upcoming peak windows are likely to be more marked than during the current flat-rate period.
Batch API Discounts
OpenAI, Anthropic, Google, and xAI all offer 50% off async batch processing with 24-hour SLAs. Groq offers 50% off batch jobs. DeepSeek's automatic prompt caching at $0.0028 per cache-hit MTok competes with formal batch discounts without requiring a separate API endpoint. Claude Sonnet 5 batch pricing is $1/$5 during the intro period vs $1.50/$7.50 standard.
Prompt Caching
Cache hit pricing across major providers:
- DeepSeek V4 Flash: $0.0028/MTok (98% off standard input)
- DeepSeek V4 Pro: $0.003625/MTok (99.2% off list price)
- Anthropic (all models): 10% of standard input ($0.20/MTok for Sonnet 5 intro, $0.50/MTok for Opus 4.8/4.7, $1.00/MTok for Fable 5)
- GLM-5.2: $0.26/MTok (81% off standard input)
- OpenAI: 10% of standard input (automatic, no setup)
- Google: 10% of standard input plus storage fees ($0.15-$1.00/1M tokens/hour)
- xAI Grok 4.5: $0.50/MTok cached (75% off), $1.00/MTok above 200K tokens
- xAI Grok 4.3/4.20: $0.20/MTok cached (~84% off)
- Groq: 50% off cached input tokens
LLM API prices moved clearly in July 2026 - Sonnet 5 intro pricing undercuts Sonnet 4.6, Fable 5 returns from suspension, and DeepSeek announces peak-hour variable rates.
Source: commons.wikimedia.org
Context Window Surcharges
Anthropic Fable 5, Opus 4.8/4.7, Sonnet 5, and Sonnet 4.6 include the full 1M context at standard rates with no surcharge. Gemini 3.1 Pro doubles input pricing above 200K tokens ($2 becomes $4, output goes $12 to $18). Gemini 2.5 Pro steps similarly ($1.25 to $2.50 above 200K). Grok 4.3 and Grok 4.20 offer 1M context at a flat $1.25/$2.50. GPT-5.6 Sol, Terra, and Luna all carry the same 1.05M context regardless of tier, with no surcharge tiering.
Grok 4.5 doubles its cached-input rate above 200K tokens ($0.50 to $1.00) despite a 500K context ceiling, and GPT-5.4 mini caps at 400K tokens. With Grok 4.1 Fast no longer pricing separately, GPT-5.6 Luna at $0.20/$1.25 with a full 1.05M context is now the most cost-efficient option in this table for long-context work without flagship pricing.
Free Tier Comparison
| Provider | Free Credits | Models Available | Rate Limits | Notes |
|---|---|---|---|---|
| Google (Gemini) | Unlimited free tier | Flash-Lite, 2.5 Flash, 3.5 Flash | 5-15 RPM, 100-1,000 RPD | Pro models paid-only |
| Groq | Free tier available | All hosted models | Varies by model | No card required |
| xAI | $175/month credits | All Grok models | Standard limits | Via data-sharing program |
| DeepSeek | 5M tokens on signup | V4 Flash, V4 Pro | Standard limits | Non-renewable |
| OpenAI | ~$5 trial credits | GPT-4o mini, limited | 3 RPM (free tier) | 3-month expiry |
| Anthropic | ~$5 trial credits | All models | Tier 1 limits | Few months expiry |
| Mistral | Free tier (limited) | Ministral 3B legacy | Rate-limited | No card required |
Google's free tier remains the most useful for development - Flash models including Gemini 3.5 Flash are free with manageable rate limits. xAI's updated data-sharing program at $175/month in credits is now the most generous paid-alternative for API access without billing commitments. The new Claude Sonnet 5 intro pricing makes Anthropic's paid tier more competitive, but there's still no free development tier comparable to Google or Groq.
Price History
Jul 15, 2026 (scheduled) - DeepSeek V4 official release; peak-hour pricing (9AM-noon and 2-6PM CST) goes live. V4 Flash peaks at $0.28/$0.56, V4 Pro at $0.87/$1.74.
Jul 9, 2026 - GPT-5.6 Sol, Terra, and Luna reach general availability at $5/$30, $2.50/$15, and $1/$6. GPT-5.4 scheduled to retire July 23.
Jul 8, 2026 - Grok 4.5 launches at $2/$6 with a 500K context window, xAI's first full-scale V9-architecture model.
Jul 1, 2026 - Claude Fable 5 API access restored at $10/$50 after a 19-day suspension under US Commerce Department export controls. Subscription caps lifted July 7.
Jun 30, 2026 - Claude Sonnet 5 launches at $2/$10 introductory pricing through August 31, 2026 (then $3/$15). Uses newer tokenizer - ~30% more tokens vs Sonnet 4.6.
Jun 30, 2026 - DeepSeek announces peak-hour variable pricing for official mid-July V4 release. Rates double during 9AM-noon and 2-6PM CST.
Jun 16, 2026 - GLM-5.2 live on standalone Z.AI API at $1.40/$4.40 with 1M context. Cache hits at $0.26/MTok.
Jun 12, 2026 - Claude Fable 5 and Mythos 5 suspended under US export-control directive. API access fully halted for both models.
Jun 1, 2026 - DeepSeek V4 Pro promotional pricing ($0.435/$0.87) becomes permanent. The original list price of $1.74/$3.48 is retired.
May 28, 2026 - Claude Opus 4.8 launches at $5/$25 standard. Fast Mode drops from $30/$150 to $10/$50 vs Opus 4.7 - a 67% reduction.
May 15, 2026 - xAI retires eight legacy Grok model slugs, including Grok 4.1 Fast. Requests silently redirect to Grok 4.3 pricing ($1.25/$2.50) rather than erroring. Full slug shutdown Aug 15, 2026.
May 2026 - DeepSeek V4 Flash arrives on the API at $0.14/$0.28. Cache hits at $0.0028.
May 19, 2026 - Gemini 3.5 Flash launches at $1.50/$9.00. Batch at $0.75/$4.50.
May 7, 2026 - Gemini 3.1 Flash-Lite moves to GA at $0.25/$1.50.
Apr 2026 - Claude Opus 4.7 launches at $5/$25. Mistral overhauls the lineup: Mistral Small 4 at $0.10/$0.30, Mistral Small 3.2 at $0.08/$0.20.
Claude Sonnet 5's tokenizer produces ~30% more tokens vs Sonnet 4.6. At the $2/MTok intro rate you still come out ahead - but after September 1, the effective rate tops Sonnet 4.6 by about 30%.
FAQ
Which LLM API is cheapest per million tokens?
Ministral 3B at $0.04/$0.04 via the legacy endpoint is the cheapest standard commercial option. For production use, DeepSeek V4 Flash at $0.14/$0.28 off-peak is far more capable. DeepSeek's peak-hour rate doubling goes live July 15, 2026.
What's the best value LLM API for production right now?
DeepSeek V4 Flash at $0.14/$0.28 off-peak with 98% cache discounts at $0.0028/MTok. For frontier-class quality at a fair price, Claude Sonnet 5 at $2/$10 intro through August 31 is the best deal in the mid-tier.
Is GPT-5.6 available yet?
Yes. GPT-5.6 Sol, Terra, and Luna went generally available July 9, 2026, at $5/$30, $2.50/$15, and $1/$6 per MTok. GPT-5.4 retires July 23 and is replaced by Terra at the same price point.
Why did Grok 4.1 Fast disappear from this table?
It doesn't have its own pricing anymore. xAI retired the slug May 15, 2026, and requests now silently redirect to Grok 4.3 at $1.25/$2.50 - more than 6x the old input rate. The slug stops resolving completely August 15, 2026.
Will DeepSeek's peak-hour pricing affect my costs?
Only during Beijing business hours (9AM-noon and 2-6PM CST) once the official V4 launch takes effect July 15, 2026. For US-based workloads those windows correspond to 1-4AM and 6-10AM Eastern, easy to avoid with async batching. Real-time stacks in Asia-Pacific timezones may see costs double.
Are there free LLM APIs for development?
Google's Gemini API has the most useful free tier - Flash models including Gemini 3.5 Flash are free with rate limits. Groq provides free LPU inference on Llama and Qwen families. xAI offers $175/month in API credits via the data-sharing program. Mistral offers rate-limited access to the legacy Ministral 3B endpoint.
Sources:
- Anthropic Claude Pricing (Official)
- Google Gemini API Pricing (Official)
- DeepSeek API Pricing (Official)
- xAI Models and Pricing (Official)
- xAI Grok Model Retirement, May 15 2026 (Official)
- Groq API Pricing (Official)
- Mistral AI API Pricing (Official)
- OpenAI API Pricing (Official)
- OpenAI API Deprecations (Official)
- Claude Sonnet 5 Launch (Anthropic)
- DeepSeek Peak-Hour Pricing (TechNode)
- GLM-5.2 API Pricing (Z.AI Official Docs)
- Claude Fable 5 Restoration (Anthropic)
- OpenAI Launches GPT-5.6 Family (TechCrunch)
- The New GPT-5.6 Family: Luna, Terra, Sol (Simon Willison)
- Grok 4.5 Pricing (eesel AI)
- LLM API Pricing Comparison July 2026 (TLDL)
✓ Last verified July 13, 2026
