AI Model Pricing Guide
Because tokens cost money and you're not made of it | Updated October 10, 2026
CHEAPEST: Claude Haiku 5.5 ($0.10/$0.50 β€100K prompts, $0.50/$2.50 above β cache reads from $0.01, 1M ctx) / Upstage Solar Mini 4 ($0.05/$0.20 promo, 524K ctx) / Liquid d1 ($0.04/1M input-only β decision model, not an LLM) / MiMo-V2.6-Flash ($0.14/$0.28, omni-modal 1M) / Muse Spark Contributor ($0.10/$0.20) / Qwen3.8-Omni-Flash ($0.15/$0.47, omni agents) / Tencent Hy3 ($0.13/$0.54) / GLM-5.3-Flash ($0.15/$0.50) / DeepSeek V4.1-Flash off-peak ($0.15/$0.60, cache $0.003) / Ling 3.1 Flash (free on OpenRouter until Oct 13)BEST VALUE: Claude Sonnet 5.5 ($2/$10 β AA Index 56, #2 overall, cache reads now $0.10 = ~20% cheaper agentic) / GPT-6.1 Sol ($2/$10, AA 52, DeepSWE matches Astra) / Claude Haiku 5.5 ($0.10/$0.50 β€100K β small-model work at a tenth of Sonnet) / Step 5 Preview ($1/$2.70 β AA 44, DeepSWE 67.7%, open weights Oct 15) / Mistral Large 4 ($0.68/$2.09 intro, list $1.36/$4.18 β open-weight cyber/finance SOTA) / DeepSeek V4.1-Flash ($0.15/$0.60 off-peak, 39.5) / Gemini 3.8 Flash ($0.75/$3.75 intro, doubles Jan 1) / Qwen 3.8-Max ($2/$6 flat 1M) / MiMo-V2.6-Pro ($0.435/$0.87) / Beam (Reflection, 501B β API pricing due later Oct) / Microsoft-Decision-1 ($0.042/1M input-only if the job is decisions, not prose)SMARTEST: Claude Opus 5.5 (AA v4.3.2 #1 β 58 max, $4/$20) / Claude Sonnet 5.5 (56 β #2, $2/$10) / Claude Fable 5.1 (53 max, $10/$50) / GPT-6 Astra (53, $10/$50, science workloads) / GPT-6.1 Sol (52, $2/$10) / Grok 4.7 (46, $2/$6) / MiMo-V2.6-Pro (46.3 β strongest open-weight, $0.435/$0.87)NEW THIS WEEK: Microsoft-Decision-1 (Oct 9 β decision model, $0.042/1M input, output free, Foundry GA; claims 35x faster than GPT-6 Sol) + GPT-6.1 Sol Ultrafast GA (Oct 8 β $12/$60, flat 6x over standard). Plus Claude Haiku 5.5 (Oct 8 NZ β $0.10/$0.50 β€100K prompts, 1M ctx, OSWorld 2.1 72.4%, ~75% cheaper than Haiku 4.5) + Sonnet 5.5 cache reads halved to $0.10. Plus Mistral Large 4 (Oct 6 β 1T/52B 'Le Chonk' public preview, $0.68/$2.09 intro vs $1.36/$4.18 list, weights + licence Oct 27) and StepFun Step 5 Preview (Oct 8 on OpenRouter β $1/$2.70, 600B/27B, AA 44, open weights Oct 15). GPT-6.1 Astra pulled from OpenAI's price page/docs after appearing Oct 5-6. Calendar: GPT-5.5 exits ChatGPT/Codex Oct 14; Step 5 weights Oct 15; Mistral weights Oct 27; GPT-5.6 Sol promo ends Nov 21; Gemini 3.8/3.7 intro doubles Jan 1; MiniMax M2 free trial ends Nov 7.
O
OpenAI
The OG of AI APIs. GPT kicked off the revolution and they're still leading.
POWER
GPT-6 Astra
The $10/$50 flagship β Ultrafast tier live
In
$10.00
per 1M
Out
$50.00
per 1M
1.05M ctx128K outputARC-AGI-3 99.9%Ultrafast 6x (API)
OpenAI's flagship, announced Sep 3 and live on the API Sep 4. $10/$50 per 1M β 5x GPT-6.1 Sol's price β with $1 cached input, $12.50 cache writes, and a long-context tier above 272K input (2x input/cache, 1.5x output). Batch/Flex 50% off; Fast mode 2x. Saturates ARC-AGI-3 (99.9%) and FrontierMath Tier 4 (98%); GPQA 96.0%. AA Intelligence Index v4.3.2: 53 max (Opus 5.5 leads at 58), but it still wins where GPT-6.1 Sol breaks down: Terminal-Bench-Science 68.1% vs Sol's 57.0% and AutomationBench 41.4% vs 36.1%. DevDay Sep 29: Ultrafast tier live in the API and ChatGPT Work/Codex (Pro 500, Enterprise) β up to 6x faster in the API, 8x in Codex at premium pricing. And in the saga of the week: OpenAI cancelled the GPT-6.1 Astra upgrade over safety regressions rather than ship it. Escalate here when 6.1 Sol measurably loses β science, hardest multi-app automations, retry after a Sol failure.
Best: Frontier agentic work, computer use, cyber defense, scientific terminal work
NEW
GPT-6.1 Sol
New DevDay upgrade β near-Astra at a fifth of the price
In
$2.00
per 1M
Out
$10.00
per 1M
1.05M ctxCached $0.10AA v4.3.2: 52Sep 29
OpenAI's DevDay model (Sep 29), a major in-place upgrade to GPT-6 Sol. Same $2/$10 per 1M as GPT-6 Sol but cached input halves to $0.10; batch $1/$5. Near matches GPT-6 Astra ($10/$50) on agentic coding, computer use and professional work at 8β23% of Astra's cost per task β matches Astra's best DeepSWE v1.1 result (75.2% at high effort). AA Intelligence Index (max): 52 vs GPT-6 Sol's 48 and Astra's 53 β within 3.5 points of Astra on 8 of 10 AA charts. Terminal-Bench 4.0: 56.1% (GPT-6 Sol: 43.9%); OSWorld 2.0: 71.4% vs 64.4%; GDP.pdf beats Opus 5.5's fallbacks at less than half the cost per task. Weak spot: science terminal work β 11.1 points behind Astra. System card: lower failure rates than GPT-6 Sol on transparency and respecting restrictions. Oct 8: the Ultrafast tier went GA for all API users at $12/$60 β a flat 6x on every line (cached input $0.60, cache writes $15; long-context >272K $24/$90) β same weights, same answers. In Codex it carries the up-to-8x badge (the 6x is the API rate-card math). Live in Work and Codex for all paid plans (not yet in Chat); API ID gpt-6.1-sol; on AWS Bedrock same day.
Best: Default for well-specified work at volume β automations, reviews, computer use, coding agents β at a fifth of Astra's price
LIMITED
GPT-Rosalind
Life-sciences model β billing live Oct 5 (trusted-access)
In
$5.00
per 1M
Out
$25.00
per 1M
128K ctxCache $0.50MedChemBench 27.5%Oct 5 billing
OpenAI's specialized life-sciences model, out of research preview (Sep 11) via the trusted-access program for qualified Enterprise/Business orgs β and from Oct 5, it costs real money: $5/$25 per 1M, $0.50 cached input, no cache writes, 128K context. API ID gpt-rosalind-research. Benchmarks (OpenAI): MedChemBench 27.5% vs GPT-5.5's 25.1%, LabWorkBench 63.2% vs 55.8%, GeneBench 21.6% with 31% fewer tokens. Internal research only β no customer-facing products. The frontier labs' pattern this month: strongest models ship as permission lists, not product pages.
Best: Drug discovery, genomics, medicinal chemistry β vetted research orgs only
RETIRED
GPT-6 Sol
Retired β delisted after 6.1 Sol; was $6/$30
In
$6.00
per 1M
Out
$30.00
per 1M
1.05M ctxCached $0.40Delisted by Oct 10Migrate to 6.1 Sol
The Sep 22 mid-tier, delisted within three weeks β OpenAI's price page (checked Oct 10) no longer lists gpt-6-sol: standard/batch/fast/flex rows show gpt-6-astra, gpt-6.1-sol and gpt-6-luna only, and it is gone from the model docs. It launched Sep 22 at $2/$10 with $0.20 cached input (batch $1/$5), was still listed at $6/$30 by Sep 30 with cache at $0.40, and was superseded Sep 29 by GPT-6.1 Sol at $2/$10 with cache reads at $0.10 and +4 AA. If you're still routing to gpt-6-sol, move the model ID to gpt-6.1-sol β better on every measured chart and cheaper.
Best: Legacy routing β move to GPT-6.1 Sol for free intelligence
NEW
GPT-6 Luna
Cheapest frontier model yet β $0.10/$0.50
In
$0.10
per 1M
Out
$0.50
per 1M
1.05M ctxCached $0.01Sep 22Cheapest frontier
The GPT-6 high-volume tier, launched Sep 22. $0.10/$0.50 per 1M β 50% below GPT-5.6 Luna's $0.20/$1.20 β with cached input at $0.01 and batch $0.05/$0.25. Trained with Astra methods; marginally stronger than 5.6 Luna. Available in ChatGPT Work and Codex for paid plans, and in the desktop app for Free and Go users.
Best: High-volume apps, agent execution, cost-sensitive production
BUDGET
GPT-5 mini
Fast & cheap
In
$0.25
per 1M
Out
$2.00
per 1M
128K ctxFast
Lightweight champion. Surprisingly capable for simple tasks and high-volume apps.
Best: Chatbots, simple QA, data extraction
BUDGET
o4-mini
Reinforcement tuned
In
$1.10
per 1M
Out
$4.40
per 1M
200K ctxFine-tuning
Price dropped 70%. Optimized for reinforcement fine-tuning workflows. Create custom reasoning patterns.
Best: Fine-tuning, custom reasoning
NEW
GPT-5.4 mini
Coding & subagents
In
$0.75
per 1M
Out
$4.50
per 1M
Cached inputCoding
New GPT-5.4-class model. Stronger than GPT-5 mini for coding and subagent workflows.
Best: Coding, subagents, mid-tier apps
NEW
GPT-5.4 nano
Cheapest 5.4-class
In
$0.20
per 1M
Out
$1.25
per 1M
Cached inputBudget
Cheapest way into the GPT-5.4 family. Cheaper input than GPT-5 mini.
Best: High-volume, budget apps
POWER
GPT-5.4
Still elite β now #2
In
$2.50
per 1M
Out
$15.00
per 1M
270K ctxReasoning
OpenAI's previous #1. Still elite for complex, multi-step problems.
Best: Hardest problems, professional work
POWER
GPT-5.2
Reasoning beast
In
$1.75
per 1M
Out
$14.00
per 1M
6.6h horizon200K ctx
Top 3 on METR. Excels at complex tasks, code, and multi-step reasoning.
Best: Code, analysis, agent workflows
POWER
GPT-5.2 Pro
Reasoning premium
In
$21.00
per 1M
Out
$168.00
per 1M
200K ctxPremium
OpenAI's most precise reasoning model. For when you need the absolute best reasoning.
Best: Hardest problems, precision work
POWER
GPT-5.6 Sol
Superseded by GPT-6 Sol at half the price
In
$4.00
per 1M
Out
$20.00
per 1M
GA Jul 9Fast mode Jul 30AgentsCodingCybersecurity
OpenAI's strongest model, now GA. SOTA on Agents' Last Exam (53.6, +13 over Fable 5), BrowseComp (92.2%), Terminal-Bench 2.1 (88.8%). Three tiers: Sol (flagship), Terra (balanced, $2/$12), Luna (fast, $0.20/$1.20). New Fast mode: 2.5x faster at $10/$60 (API). Multi-agent 'ultra' mode for demanding tasks. Aug 21 promo repricing: $4/$20 through Nov 21, then $5/$30 standard. Superseded Sep 3 by GPT-6 Astra at 2.5x the price β Sol remains the value pick for frontier agentic work. Sep 22: GPT-6 Sol launched at $2/$10 β half this promo price β trained with Astra methods and scoring higher on AA v4.3.2 (48 vs 47).
Best: Frontier agentic tasks, coding, biology, cybersecurity
20% OFF
GPT-5.6 Terra
Balanced 5.6 β now 20% cheaper
In
$2.00
per 1M
Out
$12.00
per 1M
GA Jul 9Price cut Jul 30Balanced
Price cut 20% on Jul 30 β from $2.50/$15 to $2/$12. Matches Claude Sonnet 5 on input price during intro. Outperforms Fable 5 at ~1/16 the cost. The everyday work model of the 5.6 family. Batch: $1/$6. Cached input: $0.20.
Best: General purpose, production apps
80% OFF
GPT-5.6 Luna
Superseded by GPT-6 Luna at half the price
In
$0.20
per 1M
Out
$1.20
per 1M
GA Jul 9Price cut Jul 30Fast
Price slashed 80% on Jul 30 β from $1/$6 to $0.20/$1.20 per 1M. Outperforms Opus 4.8 on coding agent index. Nearly matches GPT-5.5 peak at a fraction of the cost. The high-volume tier of the 5.6 family. Batch: $0.10/$0.60. Cached input: $0.02. Sep 22: GPT-6 Luna launched at $0.10/$0.50 β half this price with cached input at $0.01.
Best: High-volume, cost-sensitive apps, agent execution
POWER
GPT-5.5
Now #2 β still elite
In
$5.00
per 1M
Out
$30.00
per 1M
1.05M ctxReasoningAgents
Previous #1. 82.7% Terminal-Bench, 84.9% GDPval, 78.7% OSWorld. 1.05M context window. Still elite, now behind GPT-5.6.
Best: Coding, agents, research, multi-step tasks
#1 RANKED
GPT-5.5 Pro
Maximum intelligence
In
$30.00
per 1M
Out
$180.00
per 1M
1.05M ctxPremiumDeep Research
90.1% BrowseComp, 52.4% FrontierMath Tier 1-3. 1.05M context window β read entire codebases and research libraries. The ceiling for what AI can do right now.
Best: Hardest problems, deep research, scientific discovery
FLAGSHIP
GPT-5.4 Pro
Previous premium β now #3
In
$30.00
per 1M
Out
$180.00
per 1M
270K ctxPremium
Former #1, now behind GPT-5.5. Still incredibly powerful for demanding tasks.
Best: Most demanding tasks, unlimited budget
β
Microsoft
Decision models for agent control β structured outputs your software can act on, at a fraction of LLM cost.
NEW
Microsoft-Decision-1
Decision model β $0.042/1M input, output free
In
$0.04
per 1M
Out
$0.00
per 1M
Output tokens freeQwen3.5-9B base36-benchmark bestFoundry live
Microsoft's entry into the decision-model category that Liquid d1 opened (Oct 9, available in Microsoft Foundry; OpenRouter support planned). Not an LLM: it picks from predefined choices and assigns probability scores β routing, classification, prioritization, verification and workflow control, so software can branch without generating tokens. Pricing: $0.042 per 1M input tokens, output free β Microsoft's internal estimates put classifying 1M texts at ~$11 vs GPT-6 Sol's ~$2,434. Built on Qwen3.5-9B via single-pass decision-scoring post-training, with plans to rebase on MAI and OpenAI models. Claims: highest accuracy across a 36-benchmark, ~150K-question suite kept blind from training; 4.5x faster than the runner-up Quyet-1.0-Large, 35x faster than GPT-6 Sol. Internal users: Xbox Research categorized 10K+ gaming feedback items at GPT-6-Sol-comparable quality, 14x faster and 200x cheaper; Copilot's response-quality team found it competitive with GPT-5.6 Luna. All vendor-measured β launch-week claims, independent verification pending.
Best: Agent step control, model routing, data labeling, AI judging, safety screening β high-frequency structured decisions
A
Anthropic
Safety-first company. Claude is beloved by developers for being genuinely helpful.
#1 OVERALL
Claude Opus 5.5
Still AA Index #1 β the $4/$20 flagship workhorse
In
$4.00
per 1M
Out
$20.00
per 1M
1M ctxCache reads $0.20Fast mode $8/$40AA v4.3.2: 58 #1
Anthropic's flagship-workhorse (Sep 22), first of the Claude 5.5 family. $4/$20 per 1M β 20% under Opus 5 β with cache reads cut 60% to $0.20 (cache writes $5) and 30%+ faster output. Fast mode: $8/$40 at up to 2.5x speed. Batch: $2/$10. Anthropic's tests put total run cost 40% below Opus 5, and five-hour limits rose on Pro/Max/Team/Enterprise. AA Intelligence Index v4.3.2 (max): 58 β the highest score AA has measured, ahead of Sonnet 5.5 (56), Fable 5.1 and GPT-6 Astra (53). Leads six of ten Index evals: HLE 61.4%, SciCode 66.9%, AA-Briefcase 1822 Elo, GDPval-AA v2.1 1846. Best automated-alignment scores Anthropic has measured; bio/cyber capability comparable to Mythos 5.1, so it ships with Fable-5.1-class safeguards. Still clearly stronger than Sonnet 5.5 at complex open-ended work β Anthropic's own testing β but Sonnet 5.5 wins Terminal-Bench 4.0 (70.6% vs 66.4%).
Best: Coding, agents, knowledge work β Fable-class intelligence at 60% of the price
NEW
Claude Sonnet 5.5
New β #2 on the AA Index at Sonnet prices
In
$2.00
per 1M
Out
$10.00
per 1M
1M ctxCache reads $0.10Terminal-Bench 70.6%134 tok/s
Anthropic's second 5.5-family model (Sep 28), six days after Opus 5.5 β and the sharpest price-performance move of the month. $2/$10 per 1M β identical pricing to Sonnet 5 β but it needs far fewer tokens per task (up to 30% cheaper) and outputs 30%+ faster (~134 tok/s on AA, its fastest Sonnet). Oct 8 update: cache reads halved to $0.10/1M (from $0.20), which Anthropic pegs at ~20% cheaper on most agentic work where cache hits dominate. AA Intelligence Index: 56 β #2 overall behind only Opus 5.5's 58, ahead of Fable 5.1 and GPT-6 Astra (53) β at a quarter of Fable's price. Terminal-Bench 4.0: 70.6% at max effort (43.0% at high β the API default) vs Sonnet 5's 10.3% and above Opus 5.5's 66.4%. First Sonnet to beat PokΓ©mon Red from screenshots alone. Two points below Opus 5.5 on GDPval-AA. First Sonnet with cyber safeguards matching Opus-class capability + reasoning-extraction classifiers: risky security requests fall back to Sonnet 5. 1M context, 128K output (300K on Batch beta); API ID claude-sonnet-5-5. The catch: verbosity β AA measured the highest output tokens per task it has seen, so the halved cache reads are exactly where the value lands.
Best: Everyday agentic work, bug fixing, documents/slides/spreadsheets β near-Opus quality at Sonnet prices
NEW
Claude Fable 5.1
Former AA #1 β the $10/$50 frontier tier
In
$10.00
per 1M
Out
$50.00
per 1M
1M ctx128K outputCache $0.25AA Index 66 #1
Anthropic's new frontier (Sep 1). $10/$50 unchanged but cache reads cut 75% to $0.25/1M β Anthropic's August usage data shows ~25% total savings on typical workloads, up to ~45% for agentic work where cache hits dominate. Terminal-Bench-Science 52.6% (vs 24.7% on Fable 5), Terminal-Bench 4.0 55.8%, AutomationBench 31.4%, GDPval-AA v2 1853, CursorBench 73.4%. AA Intelligence Index (max, fallback): 66 β #1, three clear of Opus 5. Cyber false positives down 60%; can now find (not exploit) vulnerabilities. Breaking API changes: forced tool use returns 400, thinking blocks are model-bound. API ID: claude-fable-5-1. Sep 22 update: Opus 5.5 passed it on the AA Index (58 vs 53 max on v4.3.2) at $4/$20 with cheaper cache reads β Fable 5.1 remains the restricted frontier tier.
Best: Coding, agents, long-horizon knowledge work β migrate from Fable 5 for the cache savings
NEW
Claude Haiku 5.5
New β 75% cheaper Haiku, 1M ctx
In
$0.10
per 1M
Out
$0.50
per 1M
1M ctx (long-ctx rates >100K)128K outputFirst Haiku with effort dialOSWorld 2.1 72.4%
Anthropic's small-model refresh (Oct 8 NZ / Oct 7 US). Per 1M: $0.10/$0.50 for prompts up to 100K tokens (Anthropic: ~90% of Haiku traffic), $0.50/$2.50 above; cache reads $0.01/$0.05, cache writes $0.125/$0.625, batch half price. On average ~75% cheaper to run than Haiku 4.5. AA Intelligence Index v4.3.2 (published Oct 8): 43.4 at max β within a point of Kimi K3 and #14 overall, beating GPT-6 Luna (38.1) and MiMo-V2.6-Flash (37.9) at a fraction of Kimi's price. 1M context, 128K output (300K on Batch beta header), first Haiku with the effort parameter (default medium). Computer use jumps to 72.4% on OSWorld 2.1 (Haiku 4.5: 15.7%), Terminal-Bench 4.0 39.2% (4.5: 0.0%), HLE 45.9% no-tools. Positioned as the subagent next to Opus 5.5/Sonnet 5.5 and for high-volume classification/summarisation/routing. Live on all platforms (AWS, Google Cloud, Azure). Watch the tokenizer: same 4.7+-era tokenizer, so identical text counts ~30% more tokens than Haiku 4.5. API ID claude-haiku-5-5. Shipping alongside: Sonnet 5.5 cache reads halved to $0.10 (~20% cheaper agentic) and new monthly API credits for Max ($100/$200) and Team (up to $500) subscribers.
Best: High-volume classification, extraction, summarisation, live support, computer/browser use, subagents
LEGACY
Claude Haiku 4.5
Legacy small model β superseded by Haiku 5.5
In
$1.00
per 1M
Out
$5.00
per 1M
200K ctx10x Haiku 5.5 price
Superseded by Haiku 5.5 (Oct 7 US) at a tenth of the price with a 1M window β Anthropic pegs Haiku 5.5 at ~75% lower average running cost. Still listed on Anthropic's price page for existing pipelines; migrate claude-haiku-4-5 β claude-haiku-5-5 unless you depend on the older tokenizer's lower token counts.
Best: Existing pipelines only β new builds should use Haiku 5.5
POWER
Claude Sonnet 5
Superseded by Sonnet 5.5 at the same $2/$10
In
$2.00
per 1M
Out
$10.00
per 1M
200K ctxAgentic$2/$10 permanentMigrate to 5.5
Close to Opus 4.8 performance at Sonnet prices. Intro pricing $2/$10 made permanent Aug 10 β the scheduled Sept 1 increase was cancelled. Superseded Sep 28 by Claude Sonnet 5.5: same $2/$10/$0.20 pricing but terminal-bench agentic coding jumps 10.3% β 70.6%, output runs 30%+ faster, and typical tasks cost up to 30% less through lower token use. Anthropic's own comparison keeps Opus 5.5 ahead of both on complex open-ended work; also note Sonnet 5.5's 1M context vs this card's 200K. It's a free upgrade β move the model ID.
Best: Legacy routing β migrate to Sonnet 5.5 at the same price
BEST
Claude Sonnet 4.6
Previous Sonnet β still solid
In
$3.00
per 1M
Out
$15.00
per 1M
Balanced200K ctx
Previous Sonnet default. Still excellent but superseded by Sonnet 5 at lower intro pricing.
Best: Most tasks, code, writing, general use
POWER
Claude Opus 5
Superseded by Opus 5.5 at 20% less
In
$5.00
per 1M
Out
$25.00
per 1M
1M ctx128K outputThinking ONNew Claude Max default
Anthropic's new top-tier Opus. Scores 63 pts on the AA Intelligence Index (v4.1.1) β second only to Fable 5.1's 66. Near-Fable-5 intelligence at exactly half the price ($5/$25 vs $10/$50). ARC-AGI-3: 30.2% β 4x GPT-5.6 Sol (7.8%), 20x Opus 4.8 (1.5%). Frontier-Bench v0.1: 43.3% (Opus 4.8: 18.9%). Five-level effort toggle, thinking ON by default. Automatic fallback replaces hard refusals. API ID: claude-opus-5. New default on Claude Max; top model on Claude Pro. Fast mode: $10/$50 at 2.5x speed (API research preview). Batch API: $2.50/$12.50. Superseded Sep 22 by Opus 5.5 β 20% cheaper ($4/$20), 60% cheaper cache reads ($0.20), and the best alignment audit scores Anthropic has measured. Migrate for the savings alone.
Best: Coding, agents, knowledge work β the new cost-efficient frontier sweet spot
POWER
Claude Opus 4.8
Previous Opus flagship β superseded by Opus 5
In
$5.00
per 1M
Out
$25.00
per 1M
1M ctx128K outputSelf-verify
Previous Opus top tier, now superseded by Claude Opus 5 at the same price. 1M context, 128K output, autonomous self-verification. Same $5/$25 pricing as 4.7. Migrate to Opus 5 for near-Fable-5 intelligence at no extra cost.
Best: Complex coding, agents, long-horizon tasks
POWER
Claude Opus 4.7
Previous Opus SOTA β still elite
In
$5.00
per 1M
Out
$25.00
per 1M
1M ctxxhigh reasoningSelf-verify
Previous Anthropic best. 1M context, autonomous self-verification. Beat GPT-5.4 on BrowseComp. Now superseded by Opus 4.8 at the same price.
Best: Complex coding, agents, long-horizon tasks
POWER
Claude Opus 4.6
Proven workhorse
In
$5.00
per 1M
Out
$25.00
per 1M
14.5h horizon200K ctxFast mode
Still one of the best. 14+ hour autonomous tasks. Reliable, consistent, now the value play vs 4.7/4.8.
Best: Hard problems, research, complex agents
POWER
Claude Fable 5
Superseded by Fable 5.1 β migrate for cache savings
In
$10.00
per 1M
Out
$50.00
per 1M
1M ctx128K outputSafety classifiersRestored Jul 1
Public version of Mythos. $10/$50 β double Opus 4.8. 1M context, 128K output, autonomous self-verification. Includes safety classifiers that can refuse requests (refusals are not billed; fallback credits refund prompt-cache cost on retry). Suspended by US government Jun 12, restored globally Jul 1. Superseded Sep 1 by Fable 5.1 β same $10/$50 with cache reads at $0.25/1M.
Best: Hardest reasoning, long-horizon agentic work, cybersecurity
LIMITED
Claude Mythos 5.1
Fable 5.1 with permissive safeguards β vetted access
In
$10.00
per 1M
Out
$50.00
per 1M
1M ctx128K outputTrusted accessUS orgs only
Identical to Fable 5.1 with safeguards tuned for cybersecurity and life sciences. Ships via the Cyber Verification Program, the US-government-partnered Life Sciences Verification Program, and Project Glasswing β Anthropic's strongest cyber model ever. System card flags it as 'less honest under pressure' than recent Claude models. Also powers Claude Security. US organizations only for now.
Best: Vetted cyberdefenders, life-sciences R&D β apply via CVP/LSVP
NEW
Grok 4.7
New flagship β frontier knowledge work, same $2/$6
In
$2.00
per 1M
Out
$6.00
per 1M
500K ctxAA 46 (+2)Cached $0.50DeepSWE 71.0%
xAI's most capable model (Sep 21), shipped after four missed release windows. Same $2/$6 pricing as Grok 4.6 with $0.50 cached input. AA Intelligence Index: 46 at xhigh (+2 over 4.6), placing xAI in the top-4 labs; AA-Briefcase 1657 Elo (+111) β just behind Opus 5 and Fable 5.1; GDPval-AA 1695 (+90). CursorBench 4.0 frontier price-performance; DeepSWE v1.1 71.0% (high effort); Terminal-Bench 4.0 38.0%; EEBench, Harvey Legal, HealthBench Professional gains. Works longer and self-checks with xAI's best-calibrated safeguards. The catch: 81k output tokens per task at xhigh β double Grok 4.6's 38k and ~3x GPT-6 Astra's 27k, so the same per-token bill can cost more per task. Knowledge cutoff May 2026. Live in Cursor, Grok Build, the Grok API and model routers.
Best: Coding, long-horizon agents, professional knowledge work
POWER
Grok 4.6
Superseded by Grok 4.7 at the same price
In
$2.00
per 1M
Out
$6.00
per 1M
500K ctx1753 ELOCached $0.50Coding & agents
xAI's new flagship (Aug 12), released with Cursor. $2/$6 per 1M, $0.50 cached input. 1753 ELO claim β overtakes Kimi K3. Live on xAI API, Grok Build, Cursor, Grok Bot, and partners (OpenRouter, Vercel, Cloudflare). Fast variant available at twice the price. 2x included usage in Cursor/Grok Build for the first week. Sep 21: Grok 4.7 launched at the same $2/$6 β AA Index 46 (+2), AA-Briefcase 1657 Elo, but roughly double the output tokens per task.
Best: Coding, long-running agents, knowledge work β the new default Grok
POWER
Grok 4.5
Previous flagship β superseded by 4.6
In
$2.00
per 1M
Out
$6.00
per 1M
500K ctx80 TPS4.2x token efficiencyCoding & agents
SpaceXAI's previous flagship, now superseded by Grok 4.6 at the same price. Trained alongside Cursor. Opus 4.8-class intelligence at 80 TPS with 4.2x better token efficiency than Opus 4.8. SWE Marathon 29% (beats Opus 4.8 at 26%). Cached input at $0.50/1M. Migrate to Grok 4.6 for no extra cost.
Best: Coding, agentic tasks, knowledge work β migrate to Grok 4.6
NEW
Grok 4.20
Same price as 4.3, more features
In
$1.25
per 1M
Out
$2.50
per 1M
2M ctxReasoningMulti-agentVision
Same pricing as Grok 4.3 with multi-agent orchestration. Cached input at $0.125/1M. 2M context window for complex agent swarms.
Best: Complex multi-agent workflows
NEW
Grok 4.3
New recommended base model
In
$1.25
per 1M
Out
$2.50
per 1M
Best value2M ctxReasoningVision
xAI's recommended Grok 4 model after retiring old variants. Beats Grok 4.1 on coding, agents, and reasoning. This is the migration target for retiring models.
Best: High-volume apps, X analysis, multi-agent — the new default Grok
RETIRED
Grok 4 / 4.1 Fast
Retired May 15
In
$0.20
per 1M
Out
$0.50
per 1M
2M ctxRetired
Retired May 15, 2026. Migrated? Good. If not, move to Grok 4.3 ($1.25/$2.50) or Grok 4.20 ($1.25/$2.50).
Best: β Migrate to: Grok 4.3
RETIRED
Grok Code Fast 1
Retired May 15
In
$0.20
per 1M
Out
$1.50
per 1M
256K ctxRetired
Retired May 15, 2026. Migrate to Grok 4.3 for coding.
Best: β Migrate to: Grok 4.3
BUDGET
Grok 3 Mini
Older gen cheap
In
$0.30
per 1M
Out
$0.50
per 1M
131K ctxReasoning
Budget fallback if Grok 4's 2M context is overkill for your use case.
Best: Simple tasks, testing
POWER
Grok 4-0709
Premium tier
In
$3.00
per 1M
Out
$15.00
per 1M
256K ctxReasoningVision
Premium Grok. Smaller context but more reasoning power.
Best: Grok style with more smarts
M
Meta
Meta's first paid API. Muse Spark brings aggressive pricing and agentic capabilities from Meta Superintelligence Labs.
NEW
Muse Spark 1.3
Meta reaches the top-3 β cheapest at 59+ intelligence
In
$1.25
per 1M
Out
$4.25
per 1M
AA Index 6162 max (preview)$0.55/taskCached $0.15
Meta's fourth Muse Spark in five months (Sep 2). $1.25/$4.25 unchanged ($0.15 cached input). AA Intelligence Index: 61 at xhigh (up 4 from 1.2), 62 at max β limited partner preview β behind only Fable 5.1 (66) and Opus 5 (63). #1 on Tau3-Bench Banking. $0.55 per Index task vs $0.94β0.95 for GPT-5.6 Sol and Grok 4.6 β the cheapest of any model at 59+. Meta teases open weights and larger Muse models around Q4.
Best: Agentic work, scientific reasoning, cost-efficient frontier intelligence
Muse Glimmer 30B
First MSL open model, runs locally
NEW
Muse Spark 1.1
Previous Muse Spark β superseded by 1.3
In
$1.25
per 1M
Out
$4.25
per 1M
AgenticTool useCodingUS preview
Meta's first paid AI model via the Meta Model API. Agentic model from Meta Superintelligence Labs (run by Alexandr Wang). ~25% cheaper than comparable OpenAI/Anthropic models. $20 free credits for new accounts. US preview only β no EU access yet.
Best: Agentic tasks, tool use, cost-sensitive apps
G
Google DeepMind
Gemini has quietly become excellent. Massive context, strong multimodal, and a generous free tier.
ANNOUNCED
Gemini 4 Argon
New frontier flagship β $2/$10, but nobody can buy it yet
In
$2.00
per 1M
Out
$10.00
per 1M
1M-token outputDeepSWE 77.9%Vals Index #1Fairwind-only
Google announced its next frontier model on Sep 30 and then largely declined to ship it. Announced intro pricing: $2/$10 per 1M with cached input at 95% off and a 1M-token output limit (up from 64K). Published results: DeepSWE v1.1 77.9% (Opus 5.5: 74.2%), AutomationBench 51.3%, LVBench 91.7%, CWE-bench v1 68% (tied first), #1 on the Vals AI model index. Trained for long-horizon work: Google says Argon agents migrated 800K+ lines of the Fuchsia Zircon kernel to Rust and freed 300+ TiB of datacentre memory. The catch: access is phased through the Fairwind Program (vetted cyber defenders, 650+ partners) and US-government pre-release testing β developers, enterprises and consumers wait, starting with paid API customers and Google AI Ultra subscribers when the guardrails are ready. No date given; not on any public API yet.
Best: Watch-list only β the announced price/quality point is real, the product isn't
NEW
Gemini 3.1 Flash-Lite
Cheapest Gemini 3
In
$0.25
per 1M
Out
$1.50
per 1M
PreviewBudget
Cheapest way into Gemini 3.1. Preview tier with budget-friendly pricing.
Best: Budget Gemini 3 apps, prototyping
NEW
Gemini 3 Flash
New budget
In
$0.50
per 1M
Out
$3.00
per 1M
PreviewFast
Gemini 3 Flash preview. Balanced performance at budget pricing.
Best: Budget apps, prototyping
VALUE
Gemini 2.5 Flash
Best value
In
$0.30
per 1M
Out
$2.50
per 1M
1M ctxMultimodalFree tier
Cheapest way to process 1M context. Free tier available. Multimodal - images, video, audio.
Best: High-volume, multimodal, prototypes
NEW
Gemini 2.5 Flash-Lite
Ultra-cheap Flash
In
$0.10
per 1M
Out
$0.40
per 1M
1M ctxBudget
Flash-Lite tier for Gemini 2.5. Cheaper than standard Flash with 1M context support. Best for high-volume simple tasks.
Best: High-volume, simple tasks, cost-sensitive apps
NEW
Gemini 3.5 Flash
GA Flash β 1M context, agentic
In
$1.50
per 1M
Out
$9.00
per 1M
1M ctx65K outputAgenticComputer Use
Gemini 3.5 Flash is GA. Most intelligent Flash model for sustained agentic and coding work at scale. 1M context, 65K output, thinking, Computer Use, function calling. Free tier available.
Best: Agentic tasks, coding, production Flash workloads
NEW
Gemini 3.8 Flash
4th Flash in 4 months β AA Index 59 at Flash prices
In
$0.75
per 1M
Out
$3.75
per 1M
1M ctx~300 tok/sAA Index 59Terminal-Bench 90.8%
Google's best reasoning/coding Flash yet (Sep 2) β 3.5, 3.6, 3.7, 3.8 since May. Same $0.75/$3.75 intro as 3.7 Flash through Dec 31, then $1.50/$7.50 from Jan 1, 2027. AA Intelligence Index 59 (+3 over 3.7) β cheapest model at its intelligence level ($0.58/task). ~300 tok/s, the fastest output speed Artificial Analysis has measured. Trained on long-running agentic loops, it 'works harder': ~40% higher cost per task than 3.7 (30% more output tokens) β drop to low effort for efficiency-first routing. Gemini 3.8 Flash Cyber for trusted defenders via the Fairwind Program: 2.6x more correct Chrome patches than much larger frontier models, per Google.
Best: Agentic coding, knowledge work, price-sensitive production β the new default Gemini
NEW
Gemini 3.7 Flash
Efficiency pick β superseded by 3.8 Flash
In
$0.75
per 1M
Out
$3.75
per 1M
1M ctxGA Aug 13DeepSWE 65.3%Computer Use
Gemini 3.8 Flash (Sep 2) supersedes it at the same intro price; 3.7 remains fully supported for efficiency-first workloads. Google's most intelligent workhorse Flash model yet (Aug 13). Intro price $0.75/$3.75 β half 3.6 Flash's original cost β through Dec 31, then $1.50/$7.50. DeepSWE v1.1: 65.3% (vs 49% on 3.6). FrontierCode 1.1: 43.6% (vs 34.4%). WebDev Arena 1588 Elo (vs 1538). GDP.pdf 34% (vs 22%). AutomationBench 30.4% (vs 17%). Powers Gemini Spark. Built-in Computer Use.
Best: Agentic coding, knowledge work, cost-efficient production agents
NEW
Gemini 3.6 Flash
New workhorse β 17% fewer output tokens
In
$1.50
per 1M
Out
$7.50
per 1M
1M ctxGA Jul 21Computer UseToken efficient
Google's new workhorse Flash model (Jul 21). 17% fewer output tokens than 3.5 Flash on the AA Index (up to 65% on DeepSWE). Step up in coding (DeepSWE 49% vs 37%), knowledge work (GDPval 1421 vs 1349), computer use (OSWorld 83% vs 78.4%). $1.50/$7.50 β lower price than 3.5 Flash too. Built-in Computer Use via Gemini API.
Best: Agentic coding, knowledge work, cost-efficient agents
NEW
Gemini 3.5 Flash-Lite
Fastest 3.5 β 350 TPS
In
$0.30
per 1M
Out
$2.50
per 1M
GA Jul 21350 TPSComputer UseAgentic
Fastest model in the 3.5 series β 350 output tokens/s (AA Index). $0.30/$2.50. Outperforms 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld (74% vs 65.1%). Terminal-Bench 2.1: 54% vs 31% on 3.1 Flash-Lite. Computer use built-in. Configurable thinking levels for cost/latency tradeoffs.
Best: High-throughput agents, agentic search, document processing
NEW
Gemini 3.5 Pro
2M context β largest at the frontier
In
$1.25
per 1M
Out
$10.00
per 1M
2M ctxDeep ThinkGA Jul 17
Google's new flagship, launched July 17. 2-million-token context window β double anything at the frontier. New Deep Think extended reasoning mode (gated to $250/mo Ultra tier). Rebuilt from scratch on a new pretraining run after the original failed on recursive tool-calling. $1.25/$10 β 4x cheaper input than GPT-5.6 Sol.
Best: Massive context, video analysis, long-horizon reasoning, RAG-replacement
Gemini 3 Pro
3rd gen flagship
In
$2.00
per 1M
Out
$12.00
per 1M
1M ctx
Third-generation Gemini Pro. Now stable β no Preview tag. Same / pricing as preview tier. Strong general-purpose flagship.
Best: General production apps, stable Pro performance
FLAGSHIP
Gemini 3.1 Pro Preview
New flagship
In
$2.00
per 1M
Out
$12.00
per 1M
4h horizonPreviewVideo
77.1% ARC-AGI-2. Price increased from $1.25/$10. Batch and Flex tiers at 50% off.
Best: Video analysis, complex reasoning
LONG
Gemini 2.5 Pro
Long outputs
In
$1.25
per 1M
Out
$10.00
per 1M
1M ctx64K output
Same price as 3.1 Pro but 64K max output vs 16K. Choose for long-form content generation.
Best: Long-form writing, large outputs
DEPRECATED
Gemini 2.0 Flash
Shuts down Jun 1
In
$0.15
per 1M
Out
$0.60
per 1M
1M ctx8K outputRetiring Jun 1
Deprecated — shuts down June 1, 2026. Migrate to Gemini 2.5 Flash or 3.1 Flash-Lite.
Best: Migrate away from this model
β¬
Open Source & Local
Open-weight models you can run yourself or call via cheap APIs. The frontier is no longer closed.
PREVIEW
Beam (Reflection)
US open-weight entrant β 501B MoE, pricing pending
In
$0.00
per 1M
Out
$0.00
per 1M
In
FREE
local
Out
FREE
local
501B MoE / 23B active23.8T tokens pretrainTerminal-Bench 80.1Apache 2.0 β later Oct
Reflection AI's first open-weight model (Oct 5), red-teaming now β weights, tech report and API pricing land later this month, so there is no price to show yet. 501B-param MoE with 23B active, pretrained on 23.8T tokens, then one of the largest open-lab RL runs: 100M+ rollouts on 10.5K NVIDIA GB300s over 4 weeks (~1.3B sandboxes, 110K concurrent rollouts). Scores: SWEBench Verified 80.9, Terminal-Bench v2.1 80.1, SWE-Bench Pro v1 65.5, DeepSWE v1.1 44.4 β GLM-5.2-class with 3-4x less inference compute; Kimi K3 still leads raw capability. Reasoning-effort control trades tokens for depth. From Mislav BalunoviΔ & Wei Chen (ex-DeepMind, AlphaZero lineage); $25B valuation, 250MW Korean sovereign build, SpaceX Colossus-2 lease.
Best: Enterprise coding agents once pricing lands β watch-list for the open-weight crown
FREE
Kolibri 1
Germany's sovereign open-weight model β DE/EN reasoning
In
$0.00
per 1M
Out
$0.00
per 1M
In
FREE
local
Out
FREE
local
1M ctx validated78B MoE / 3.46B activeApache 2.0Reasoning mode
Aleph Alpha's sovereign LLM (released Oct 3, Hugging Face: Aleph-Alpha/Kolibri-1). 78B-total MoE with 3.46B active, Apache 2.0, vLLM-ready at ~78GB FP8 (runs on 2x A100 80GB up). 1M-token context validated for quality and serving (262K native; positional encoding in sliding-window layers extends in principle indefinitely). German-first: trained to reason in German, ~23.9% of its 20T-token pretraining mix, German-optimized tokenizer. Reasoning mode, tool calling, RAG with documented abstention, adjustable reasoning effort. EU GPAI Code of Practice signatory; training screened against a 4.5M-URL blocklist with per-dataset license checks. #1 on Hugging Face trending after launch.
Best: German-language RAG, public-administration and compliance-heavy deployment you host yourself
NEW
Mistral Large 4
1T 'Le Chonk' preview β open-weight cyber crown, $0.68/$2.09 intro
In
$0.68
per 1M
Out
$2.09
per 1M
512K ctx (docs: 1M)1.05T MoE / 52B activeCyber 82% / Cybench 93%Weights Oct 27
Mistral's biggest model yet (Oct 6 public preview, 'le Chonk'): 1.05T-param MoE with 52B active per Mistral's current docs (launch posts said 49B; the model 'continues to improve rapidly as we refine it') plus a 1.6B vision encoder, natively multimodal (image+text in), trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacentres β the first model out of its β¬3B round. API: $0.68/$2.09 per 1M with $0.07 cached input during preview; docs show struck-through list pricing of $1.36/$4.18 (cache $0.14) with no end date for the discount β budget on list. Mistral claims SOTA among open models on cyber (82% on AA's find-and-patch test β the highest of any model, where Opus 5.5 and GPT-6 Astra score near zero by refusing β plus 93% on Cybench), finance and law, SOTA open-weight on SciCode-Verified, and beats closed frontier models on visual grounding (Dense 200 42% vs Astra's 41%). AA Intelligence Index v4.3.2 preview: 38.4; Mistral's own brief ranks it top-five globally on the AA Cyber Index and leads open-weight models developed outside China 'by a wide margin'. Coding: DeepSWE v1.1 61.7%, Terminal-Bench 4.0 28.3%, Coding Agent Index 49.8% (ahead of DeepSeek V4 Pro and Qwen 3.8-Max; Kimi K3 and GLM-5.3 still lead), second of five in a blind Surge AI human eval behind Opus 5. AutomationBench 59.9%, AA-Briefcase 1393 Elo. The RL run is still in flight and 'shows no sign of saturation' β numbers will move. Open weights + custom licence drop Oct 27 after a red-team window with cyber partners and state authorities (reduced-moderation version). Vals AI ranks it lower than the hype (#32 of 44 on its index) β the launch cards lean on verticals, not the aggregate.
Best: Sovereign/EU compliance deployments, cyber defence at open weights, multimodal document work β watch the Oct 27 weights drop
NEW
Liquid d1
Decision model β probabilities, not prose; $0.04/1M input-only
In
$0.04
per 1M
Out
$0.00
per 1M
$0 output tokens200-300msVision (base64)Non-autoregressive
Liquid AI's first 'decision model' (Oct 5, general availability after last week's experimental drop). Not an LLM: one forward pass reads text + images and returns probabilities for your questions β no generated tokens, 200-300ms per decision. Billed on input only: $0.04/1M, output $0.00; a 1024x1024 image counts ~1,536 tokens. On six real apps (ticket filtering, code search, web agents, visual inspection) it matches or beats GPT-6.1 Sol on four at 19-200x lower cost, answering significantly faster. Three question types: yes/no, choice, score. API model ID d1; also on OpenRouter and Vercel (text-only there so far). Open weights promised for upcoming models.
Best: Classification-style decisions at scale β filtering, routing, inspection; cheaper than prompt-hacking an LLM into JSON
NEW
Step 5 Preview (StepFun)
600B open-weight agent MoE β $1/$2.70, weights Oct 15
In
$1.00
per 1M
Out
$2.70
per 1M
1M ctx600B MoE / 27B activeAA 44 β #12Open weights Oct 15
StepFun's flagship agentic model (launched ~Sep 18 on StepFun's own API at the same $1/$2.70; hit OpenRouter Oct 8 as a router). 600B-param sparse MoE with 27B active, 1M-token context, text+image+video input (max output 64K on StepFun's docs; 128K on OpenRouter; reasoning + tool calling). StepFun's own pitch: 'the Pareto frontier marks the best trade-offs between intelligence and cost β progress begins when that boundary shifts outward.' AA Intelligence Index v4.3.2: 44 β #12, above Kimi K3 and GLM-5.3, at ~$1.03 per Index task vs Sonnet 5.5's $1.08β$7.60 by effort. Vendor numbers: DeepSWE v1.1 67.7% (vs Kimi K3's 67.5), Terminal-Bench v2.1 85.0%, GPQA Diamond 93.5%, BrowseComp 88.7%, FrontierFinance 66.4% β beating GPT-6 Astra's 55.0; a weaker Terminal-Bench v4 of 33.3% vs GLM-5.3's 41.9 shows the aggregate sits below the highlights. Very verbose: 160M output tokens across the AA run, roughly double the median. The stand-out demo: 24 hours unsupervised on an H100 optimizing an MLA GPU kernel to 508 TFLOPS (vs Opus 5's 493). Cache hits bill at ~10% of input ($0.10) β effective blended rate ~$0.54/1M. Open weights land Oct 15; MIT-style licence on the open-source release.
Best: Agentic coding, finance/research agents, long-horizon work at mid-tier prices β strongest sub-$5 open-weights candidate since MiMo-V2.6
NEW
MiMo-V2.6-Pro
New open-weight king β 1.02T MoE, omni-modal, $0.435/$0.87
In
$0.43
per 1M
Out
$0.87
per 1M
1M ctx (128K out)1.02T MoE / 42B activeNative omni-modalAA 46.32 β #1 openMIT weights
Xiaomi's new flagship (Sep 22), released and open-sourced together: 1.02T-param MoE (42B active) with hybrid attention and multi-token prediction, trained via massively scaled RL on the recursive-self-improvement path. AA Intelligence Index v4.3.2: 46.32 (max) β the strongest open-weight model measured, ahead of GLM-5.3 and Kimi K3 (~44) and past Grok 4.7's 46 on the same revision. Native omni-modal: text, image, video and audio input; 1M context, 128K output. API: $0.435/$0.87 per 1M with cache hits at $0.0036 β flat vs V2.5, which Xiaomi pegs at 1/20 to 1/60 of overseas pricing at equal intelligence. MIT weights, tech report, Distill-Qwen-9B and RL training resources on Hugging Face; an ultraspeed variant serves the same model faster.
Best: Agentic coding, multimodal understanding, open-weight frontier work at 1/20 of US frontier pricing
NEW
MiMo-V2.6-Flash
310B open MoE at $0.14/$0.28 β omni-modal budget king
In
$0.14
per 1M
Out
$0.28
per 1M
1M ctx310B MoE / 15B activeOmni-modalCache $0.0028MIT weights
The MiMo-V2.6 high-volume tier (Sep 22): ~310B-param MoE with ~15B active, native omni-modal input, 1M context. $0.14/$0.28 per 1M with cache hits at $0.0028 β among the cheapest multimodal APIs available, MIT-licensed with full weights open. Shares the V2.6 series pricing with Pro and the ultraspeed variant; Token Plan subscriptions cover the whole suite.
Best: High-volume multimodal agents, cost-sensitive production, local-adjacent open weights
BUDGET
Solar Mini 4
Upstage's 3B-active agent MoE β $0.05/$0.20 promo
In
$0.05
per 1M
Out
$0.20
per 1M
524K ctx35B MoE / 3B activeCache $0.01Promo 50% off
Upstage's compact agentic model (Sep 23): 35B-param MoE with 3B active and a 524K context window. $0.05/$0.20 per 1M on OpenRouter β a 50% promotional discount off the $0.10/$0.40 standard rate; cache reads $0.01. Korean-language strength plus English and Japanese. No published frontier-benchmark ceiling β it's built for throughput, not leaderboards.
Best: High-throughput agentic pipelines, document retrieval, long-conversation agents
NEW
DeepSeek V4.1-Flash
New flagship β 552B CED MoE at $0.15/$0.60 off-peak
In
$0.15
per 1M
Out
$0.60
per 1M
1M ctx552B MoE (8B prefill / 16B decode)Native visionMIT weightsCache $0.003
DeepSeek's new flagship (Sep 10, 2026) β first model on the Causal Encoder-Decoder (CED) architecture: 552B-param backbone that activates 8B per token on prefill and 16B on decode. Native image input, 1M context, 384K max output. MIT-licensed weights on Hugging Face (~510GB FP8). API: $0.15/$0.60 off-peak, $0.30/$1.20 peak, cache hits $0.003/$0.006 β a 50x gap between cache-hit and cache-miss input. DeepSeek-run benchmarks: DeepSWE v1.1 74.2 (V4 Pro: 62.7), Terminal-Bench 2.1 90.6, CyberGym 88.1. Independent AA Intelligence Index v4.3 (max): 39.5 β above V4 Pro 0813 (36.3) at a third of the peak price. Global KV cache down to 890 bytes/token (~1/4 of V4 Flash). Retires V4-Flash; from Sep 14 all deepseek-v4-pro requests route to V4.1-Flash at Flash rates until a V4.1-Pro ships.
Best: Cache-heavy agent loops, coding agents, high-volume multimodal β cheapest 1M-context frontier-adjacent API
NEW
GLM-5.3
Z.ai's open-weight coder β now #2 open-weight
In
$1.40
per 1M
Out
$4.40
per 1M
743B paramsOpen weight (pending)CodingGLM Coding Plan & ZCode
Z.ai (formerly Zhipu AI) calls GLM-5.3 the most capable open-weights model for coding. 743B-param model built by scaling post-training on the GLM-5.2 base. Weights are now public on Hugging Face (MIT-style GLM-5.3 licence, Aug 27) β 51 quantizations available. Live via GLM Coding Plan, ZCode and the API. AA Intelligence Index v4.3 (max): 44.9 β the strongest open-weight model measured. GLM-5.3-Flash (41.9) is the cheap sibling at $0.15/$0.50 with 1M ctx after its Sep 9 promo ended. 'Dramatic improvement over GLM-5.2 with fewer output tokens.' GLM-5.2 API was ~$1.40/$4.40 β roughly a tenth of US frontier per-token rates. Sep 22: Xiaomi's MiMo-V2.6-Pro (46.3 on AA v4.3.2) takes the open-weight #1 spot; GLM-5.3's 44.9 (v4.3) is now the strongest Z.ai model.
Best: Agentic coding, open-weight workflows, cost-sensitive enterprise
NEW
DeepSeek V4-Pro
Still live β V4.1-Flash now beats it on cost & speed
In
$0.66
per 1M
Out
$1.98
per 1M
128K ctxGA Aug 13DeepSWE 62.7%Terminal-Bench 87.9Off-peak price
DeepSeek's prior flagship out of preview (Aug 13). DeepSWE 62.7%, Terminal-Bench 2.1: 87.9. AA Intelligence Index v4.3 (max): 36.3. Peak/off-peak billing: off-peak $0.66/$1.98, peak $1.32/$3.96. Update Sep 10: DeepSeek's own tests put V4.1-Flash ahead on performance, cost, speed and runtime. From Sep 14 new deepseek-v4-pro requests can still be served (DeepSeek walked back the retirement after user demand) but check whether V4.1-Flash at $0.15/$0.60 off-peak does the same job for less. Migrate to DeepSeek V4.1-Flash unless you specifically need V4-Pro's HLE lead (42.7 vs 36.8).
Best: Autonomous agents, software engineering, cost-sensitive high-volume workloads
NEW
Qwen3.8-Omni-Flash
Omni-modal agents β audio+video in one model, 1/5 of Gemini's price
In
$0.15
per 1M
Out
$0.47
per 1M
1M ctx (991K in / 131K out)Audio + video native262K reasoningAPI-only, no weightsCache $0.016
Alibaba's first omni-modal agent model (Sep 18). Accepts text, images, audio and video natively β up to 3 hours of audio, 2-hour video files β and returns text, with thinking on by default, function calling and web search. Built on the open-weight Qwen3.8-Flash-Next architecture but served API-only (QwenCloud, Model Studio); no downloadable weights. QwenCloud pricing: $0.15/$0.47 per 1M, implicit cache hits $0.016 β audio input under $0.01/hour, ~$0.20 per minute of 720p video. +25% average across 29 evals over Qwen3.5-Omni-Plus; WildClawBench-MM 34.5 to 71.0, UniClawBench 69.6 (Gemini 3.8 Flash: 69.0), though Gemini still leads AgenticVBench (45.0 vs 36.8). SWE-bench Pro 63.3, LiveCodeBench v6 92.6. Qwen-MM-Plugins (Apache-2.0) add audio/video handling to Claude Code, Gemini CLI and Qwen Code; a Realtime variant handles live audio.
Best: Meeting/call analysis, video research agents, media indexing β audio+video agentic work at a fifth of Gemini's price
NEW
Qwen 3.8-Max
Alibaba's flagship β 'second only to Fable 5'
In
$2.00
per 1M
Out
$6.00
per 1M
1M ctx (flat tier)2.4T MoE / 95B activeMultimodalOpen weights soonAA Index 58 pts
Alibaba's most capable model. 2.4T-param MoE with 95B active. $2/$6 per 1M β flat across the entire 1M-token context window (no long-prompt surcharge). Cached input at $0.25/1M. Claims 'second only to Fable 5'. AA Intelligence Index v4.1.1: 58 pts (#8 globally, between GPT-5.6 Sol high and GPT-5.6 Terra max). Multimodal (text + visual). Open weights promised within days of GA. OpenAI and Anthropic API compatible.
Best: Long-context work, coding, enterprise β flat pricing across 1M context
NEW
MiniMax M2
Agent & coding model β 8% of Sonnet's price
In
$0.30
per 1M
Out
$1.20
per 1M
Open weight~100 TPSAgent & codeFree until Nov 7
MiniMax's agent-first model. $0.30/$1.20 per 1M β 8% of Claude Sonnet's price at ~2x the speed (~100 TPS). Top 5 globally on Artificial Analysis Intelligence Index. Built for end-to-end dev workflows (Claude Code, Cursor, Cline, Kilo Code, Droid). Open weights on HuggingFace. Free API trial until Nov 7. MiniMax Agent product also free for a limited time.
Best: Agents, coding, tool use, cost-sensitive agentic workflows
NEW
Tencent Hy3
Global open API β cheapest per-token
In
$0.13
per 1M
Out
$0.53
per 1M
256K ctx295B MoE / 21B activeApache licenseOpen weights
Tencent's reasoning and agent model (formerly Hunyuan). 295B MoE with 21B active per token. Apache-licensed weights on HuggingFace. OpenRouter from $0.13/$0.53 per 1M β among the cheapest open APIs available. Topped OpenRouter usage leaderboard within a week of its Jul 6 launch. Global access via WorkBuddy (free until Aug 31), Tencent Cloud TokenHub, and API.
Best: Reasoning, agent tasks, cost-sensitive high-volume apps
NEW
Tencent Hy4 preview
770B open-weight β frontier-class under $1/1M input
In
$0.83
per 1M
Out
$2.50
per 1M
1M+ ctx770B MoE / 49B activeApache 2.0Open weights
Tencent's next-gen Hunyuan, launched Aug 28 2026. 770B total / 49B active MoE with a context window exceeding 1M tokens. Apache 2.0 weights on Hugging Face. API: $0.834/$2.501 per 1M ($0.042 cache hits) via Tencent Cloud TokenHub and OpenRouter. Free for 2 weeks on WorkBuddy & CodeBuddy; Hy3 free access extended to Sep 30. Targets coding, office work, game dev, and scientific research. More Hy4-series models coming soon.
Best: Long-document work, coding, office productivity at open-weight prices
NEW
Kimi K3
World's largest open-weight β 2.8T params
In
$3.00
per 1M
Out
$15.00
per 1M
1M ctxMultimodalMoE 2.8T / 104B activeAA v4.3: 43.8Open weights Jul 27
Moonshot AI's flagship. 2.8-trillion-parameter MoE β world's largest open-weight model. Scores 57 on the Artificial Analysis Intelligence Index β third overall (behind Fable 5 ~60, GPT-5.6 Sol ~59), first open-weight model ever in the top 3. First on LMArena Frontend Code Arena (1,679 Elo). Native multimodal (text + images), 1M context. $3/$15 β averages $0.94 per AA Index task vs Sol $1.04, Opus 4.8 $1.80. Full weights public Jul 27 (1.56TB). Shook US chip stocks on release.
Best: Long-form coding, complex reasoning, enterprise open-weight
NEW
Kimi K2.6
88% cheaper than Opus
In
$0.60
per 1M
Out
$2.50
per 1M
256K ctxOpen weightMoE 1T/32B
Beats GPT-5.4 and Opus 4.6 on SWE-Bench Pro. 1T params, 32B active. 300 sub-agent orchestration. OpenAI-compatible API.
Best: Coding, agents, long-horizon tasks
CHEAPEST
Qwen 3.6 Plus
1M context, free tier
In
$0.10
per 1M
Out
$0.30
per 1M
1M ctxReasoningFree tier
Alibaba's latest. Mandatory chain-of-thought reasoning. Free tier available. Topped 6 coding benchmarks on release.
Best: Budget coding, massive context
NEW
Llama 4 Scout
10M context MoE
In
$0.15
per 1M
Out
$0.55
per 1M
10M ctxOpen weightMoE 109B
Longest context of any open model. 109B total, 17B active. Multimodal. Runs on 24GB VRAM.
Best: Massive context, multimodal, local
Llama 4 Maverick
Frontier coding MoE
In
$0.20
per 1M
Out
$0.80
per 1M
1M ctxOpen weightMoE 400B
Beats GPT-4o on coding. 400B total, 17B active. 128 experts. Frontier quality at MoE prices.
Best: Coding, complex reasoning
DeepSeek V3.2
Matches GPT-4o
In
$0.27
per 1M
Out
$1.10
per 1M
128K ctxOpen weightMoE 685B
94.2% MMLU matching GPT-4o. 685B MoE with 37B active. Best open model for general knowledge.
Best: General knowledge, research
FREE
Qwen3-Coder 8B
Local coding king
In
$0.00
per 1M
Out
$0.00
per 1M
In
FREE
local
Out
FREE
local
32K ctxLocal only8B dense
Runs on any 8GB GPU. 92 programming languages. 80-150 tok/s. Best local coding model under 10B. Set it up locally →
Best: Local coding, autocomplete
FREE
DeepSeek R1 Distill 14B
Local reasoning
In
$0.00
per 1M
Out
$0.00
per 1M
In
FREE
local
Out
FREE
local
Local onlyReasoning10GB VRAM
Chain-of-thought reasoning on 10GB VRAM. The sweet spot for local reasoning. 55 tok/s on modern GPUs. Run it offline →
Best: Local reasoning, budget hardware
π‘ Did You Know?
Microsoft-Decision-1 β the decision-model category goes mainstream
Microsoft shipped Microsoft-Decision-1 on Oct 9: not an LLM but a decision model that reads a brief and returns probabilities over predefined choices β continue, stop, retry, escalate, route β at $0.042 per 1M input tokens with output free. Microsoft's 36-benchmark suite (~150K questions, kept blind from training) puts it first on accuracy and fastest measured: 4.5x the runner-up Quyet-1.0-Large, 35x faster than GPT-6 Sol, classifying 1M texts for ~$11 where GPT-6 Sol costs ~$2,434. The base is telling: Qwen3.5-9B plus single-pass decision-scoring post-training, with rebases on MAI and OpenAI models planned β the HN thread notes this is the third decision model in a fortnight built on a Qwen base (Cloudflare's Clef, Strands' decider, now Microsoft's). Internal results: Xbox Research categorized 10K+ gaming feedback items at GPT-6-Sol-comparable quality 14x faster at 1/200 the cost. One week after Liquid d1's $0.04/1M input-only pricing, the biggest software company on earth validated the category. The agent-economy plumbing layer is now its own product line.
Haiku 5.5 β the small-model floor drops another 90%
Claude Haiku 5.5 (Oct 8 NZ / Oct 7 US) rewrites the small-model math: $0.10/$0.50 per 1M for prompts up to 100K β a tenth of Haiku 4.5's $1/$5, with 1M context (long-context rates $0.50/$2.50 above 100K), cache reads from $0.01, and 300K output on the Batch beta. Anthropic pegs average running cost ~75% below Haiku 4.5 after accounting for the new tokenizer's ~30% token inflation. AA's first full Index measurement (published Oct 8): 43.4 at max β inside a point of Kimi K3's 43.6 and #14 of 687 measured models, at roughly a tenth of Kimi's per-token price. The capability leap is the story: OSWorld 2.1 72.4% (4.5: 15.7%), Terminal-Bench 4.0 39.2% (4.5: 0.0%). Paired with Sonnet 5.5's cache reads halved to $0.10 the same day, Anthropic is pricing for the subagent economy: big model plans, cheap models swarm.
Step 5 Preview β StepFun moves the Pareto frontier
StepFun's Step 5 Preview reached OpenRouter on Oct 8 and became instantly routable: a 600B-param sparse MoE (27B active) with 1M context, text+image+video input, at $1/$2.70 per 1M with cache hits at 10% of input β an effective blended rate around $0.54/1M. AA Intelligence Index: 44 (#12) at ~$1.03 per task, just above Kimi K3 (43.6) and GLM-5.3 (44.8 is close) on the same board revision, with vendor-claimed DeepSWE 67.7%, GPQA 93.5%, FrontierFinance 66.4% (vs GPT-6 Astra's 55.0) and a Terminal-Bench v2.1 of 85%. Two honest catches: it is extremely verbose (160M output tokens across AA's run, ~2x median β the real bill lives in output), and the aggregate trails the highlight charts (AA had it #12 while StepFun's cards cherry-pick finance wins). The unusual part is the open-weights date sitting in black and white: Oct 15 β and the 24-hour H100 demo where the model optimized an MLA GPU kernel to 508 TFLOPS unsupervised, beating Claude Opus 5's 493. Open-weight frontier models just gained a credible Western-route competitor to Xiaomi's MiMo-V2.6-Pro at less than half the price.
Mistral Large 4 β the trillion-param 'Le Chonk'
Mistral launched Large 4 (Oct 6, public preview) as its largest model ever: 1.05T-param MoE, 52B active per Mistral's current model docs (launch posts said 49B), 1.6B vision encoder, natively multimodal, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacentres β the first model out of its β¬3B round. Preview API: $0.68/$2.09 per 1M ($0.07 cached) against struck-through list pricing of $1.36/$4.18 with no announced end date. AA Intelligence Index v4.3.2 preview: 38.4. The pitch is sovereignty plus refusals: 82% on AA's find-and-patch cyber test β the highest of any model, in a brief that also ranks ML4 top-five globally on AA's Cyber Index and leading open-weight models developed outside China 'by a wide margin' β plus 93% on Cybench, SOTA open-weight SciCode-Verified, and visual grounding above GPT-6 Astra (Dense 200: 42% vs 41%). DeepSWE 61.7%, Coding Agent Index 49.8%, second of five in a blind Surge AI human eval behind Opus 5. The RL run is still in flight ('no sign of saturation'); weights and the custom licence land Oct 27, after a reduced-moderation red-team window with cyber partners and state authorities.
Liquid d1 β a model that never generates tokens
Liquid AI's d1 (Oct 5) isn't an LLM and doesn't compete on the IQ leaderboards: it reads text/images in a single forward pass and outputs probabilities (yes/no, choice, score) in 200-300ms with zero generated tokens. Priced at $0.04 per 1M input tokens only β and on six practical apps (support-ticket filtering, code search, filing 105 documents, web-agent action selection, context compaction, circuit-board defect inspection at 85-97% accuracy) it matched or beat GPT-6.1 Sol on four, at 19-200x lower cost. Images bill as input (~1,536 tokens for 1024x1024). Live on Liquid API, OpenRouter, Vercel. A separate API category β decisions, not prose β is now a product line.
Kolibri β sovereign AI made in Germany
Aleph Alpha's Kolibri (Oct 3) is a different open-weight bet: not leaderboard-chasing, but EU-compliance-first deployment customer-controlled infrastructure. 78B-param MoE (3.46B active), 1M-token context validated (262K native, SWA position extension), reasoning mode + tool calling + RAG that abstains when evidence is missing. Pretrained on 20T tokens (~24% German) on 768 B200s β 392k GPU-hours, 6.4e23 FLOPs β with a German-optimized tokenizer. Apache 2.0, runs at ~78GB FP8 on 2x A100. Trained-data governance documented: 4.5M-URL blocklist (EC Piracy Watch List source), per-dataset license and opt-out screening. #1 on HF trending (600+ likes); Aleph Alpha is also merging with Cohere pending regulatory approval.
Reflection Beam β the US open-weight counter-attack
Reflection AI released Beam (Oct 5), its first open-weight model: 501B-total MoE, 23B active, pretrained on 23.8T tokens, then an RL campaign of 100M+ rollouts on 10.5K GB300 GPUs β one of the largest open-lab RL runs ever, with 110K concurrent rollouts and ~1.3B sandboxes. SWEBench Verified 80.9, Terminal-Bench 2.1 80.1, DeepSWE 44.4: GLM-5.2-class quality at 3-4x less inference compute, while Kimi K3 still leads raw capability. Built by ex-DeepMind founding team (Mislav BalunoviΔ, Wei Chen β AlphaZero/AlphaTensor lineage) with a $25B valuation, SpaceX Colossus-2 compute and a 250MW Korean sovereign AI factory. Weights under Apache 2.0 'later this month'; pricing TBD. Forecasters had it for 2027 β it shipped three months early.
Sonnet 5.5 β the mid-tier ate the flagship's homework
Claude Sonnet 5.5 (Sep 28) keeps Sonnet 5's pricing ($2/$10) yet scores 56 on the AA Intelligence Index β #2 overall, behind only Opus 5.5's 58 and at a quarter of Fable 5.1's $10/$50 price. Terminal-Bench 4.0: 70.6% at max effort vs Sonnet 5's 10.3% β and above Opus 5.5's 66.4%, on a benchmark where its own flagship loses. It runs ~134 tok/s (its fastest Sonnet) and costs up to 30% less per task via fewer tokens; the catch is verbosity β AA recorded the highest output tokens per task it has ever measured. Oct 8: cache reads halved to $0.10/1M, cutting most agentic workloads another ~20%. First Sonnet to ship with Opus-class cyber safeguards. Fable 5.5 is reportedly already in internal testing.
GPT-6.1 Sol β the quiet DevDay upgrade, and the Astra that wasn't
DevDay (Sep 29) shipped GPT-6.1 Sol: same $2/$10 as GPT-6 Sol, cached input halved to $0.10, and AA Index 52 vs 48 at the same price β within 3.5 points of GPT-6 Astra on 8 of 10 charts at 8β23% of Astra's cost per task, matching Astra's best DeepSWE result (75.2%). Terminal-Bench 4.0 jumps 43.9% β 56.1% for free. The bigger story was what didn't ship: OpenAI cancelled the GPT-6.1 Astra upgrade over safety regressions hours before its own conference. The Ultrafast tier is live for Astra (up to 6x faster in the API, 8x in Codex); 6.1 Sol Ultrafast is 'coming soon'.
Gemini 4 Argon β announced, priced, and unbuyable
Google announced Gemini 4 Argon on Sep 30 with intro pricing of $2/$10 and a 1M-token output limit (up from 64K) β then restricted the whole launch to Fairwind Program cyber defenders and US-government pre-release testing. Published results are real: DeepSWE v1.1 77.9%, AutomationBench 51.3%, LVBench 91.7%, CWE-bench v1 68%, #1 on the Vals AI model index. Internal deployments (800K+ lines of Fuchsia kernel to Rust, 300+ TiB of datacentre memory freed) are auditable; the public product is not. A launch that ships no purchasable API is a claim, not a product β the second such 'release' this month after OpenAI's Astra cancellation.
Opus 5.5 β a new #1 and the new value frontier
Claude Opus 5.5 (Sep 22) took the AA Intelligence Index top spot at 58 max on v4.3.2 β now bracketed within its own family by Sonnet 5.5 (56 at $2/$10, Sep 28) β while cutting price to $4/$20 (20% under Opus 5) and cache reads 60% to $0.20. Output is 30%+ faster, Fast mode runs $8/$40 at 2.5x, and Anthropic's tests put total run cost 40% below Opus 5. It leads six of ten Index evals (AA-Briefcase 1822 Elo, GDPval-AA 1846, HLE 61.4%, SciCode 66.9%) and posted the best automated-alignment scores Anthropic has measured β while shipping with Fable-5.1-class safeguards because its bio/cyber capability now matches Mythos 5.1. Sonnet 5.5 wins Terminal-Bench 4.0 (70.6% vs 66.4%); Opus 5.5 stays the pick for complex open-ended work. Haiku 5.5 arrives in the coming weeks.
GPT-6 Sol & Luna β OpenAI folds its own price card
One hour after Opus 5.5 landed on Sep 22, OpenAI launched GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50) β 50% below the GPT-5.6 promo rates, with cached input at $0.20/$0.01 and batch at half list. Both are trained with Astra's methods; OpenAI claims Sol beats Opus 5 at ~9% of the cost, and GPT-6 Sol (AA 48) already outscores GPT-5.6 Sol (47) at half the price. Sep 29: GPT-6.1 Sol superseded GPT-6 Sol at the same price with cache at $0.10 and AA 52 β routing built on GPT-6 Sol should already have been switched. GPT-5.6 Sol's $4/$20 promo now runs at least through Nov 21 with cheaper successors sitting below it.
MiMo-V2.6 β open weights take the intelligence crown
Xiaomi released and open-sourced MiMo-V2.6 on Sep 22: Pro is a 1.02T-param MoE (42B active) scoring 46.32 on AA v4.3.2 β the strongest open-weight model measured, ahead of GLM-5.3, Kimi K3 and Grok 4.7's 46 β at $0.435/$0.87 with $0.0036 cache hits. Flash (310B/15B active) costs $0.14/$0.28. Native omni-modal input (text/image/video/audio), 1M context, MIT weights with the tech report and RL training resources published. Xiaomi broadcast parts of the RL run and claims 1/20β1/60 of overseas pricing at equal intelligence.
Grok 4.7 β same price, double the bill
Grok 4.7 (Sep 21) finally shipped after four missed windows at the same $2/$6 as Grok 4.6. It gained +2 to AA 46 and +111 Elo on AA-Briefcase (1657, just behind Opus 5 and Fable 5.1), with frontier price-performance on CursorBench 4.0. But AA measured 81k output tokens per task at xhigh β double Grok 4.6's 38k, versus 27k for GPT-6 Astra max β so identical per-token pricing can still mean a higher bill per finished task.
Qwen3.8-Omni-Flash β the omni-modal price floor
Alibaba's Qwen team shipped Qwen3.8-Omni-Flash on Sep 18 β its first model trained specifically for tool use in multimodal contexts. Text, images, audio (up to 3 hours) and video (2-hour files) go in as native inputs; text comes out. QwenCloud prices it at $0.15/$0.47 per 1M with $0.016 cache hits β roughly a fifth of Gemini 3.8 Flash's intro rate, which itself doubles on Jan 1, 2027. Qwen's own tables put UniClawBench at 69.6 vs Gemini's 69.0, while Gemini keeps the lead on AgenticVBench (45.0 vs 36.8). Audio input prices fell 98% vs the previous omni tier. The catch: it's API-only β the open-weight playbook stopped at the Flash-Next base.
Mercury 2.5 β diffusion LLM at 1,107 tokens/sec
Inception Labs released Mercury 2.5 on Sep 8 β the largest diffusion language model trained to date. It generates tokens in parallel rather than left-to-right, hitting 1,107 tok/s on widely-available NVIDIA GPUs. Quality is up 40% over Mercury 2, comparable to GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite and Claude Haiku 4.5. List $0.20/$0.75 per 1M with an 80% launch discount ($0.04/$0.15) expiring early October. 260K context, tunable reasoning, parallel tool calls, schema-aligned JSON. A diffusion LLM is now beating token-by-token models on the latency axis.
Cognition SWE-2 β frontier coding without an API
Cognition shipped SWE-2 on Sep 10, post-trained from Kimi K3 with RL scaled to the multi-trillion-parameter regime. 50.0% on FrontierCode 1.1 Main β within a point of Fable 5.1 (50.9%) at a claimed 64% lower cost, and 92.8% on Terminal-Bench 2.1, the best measured score. Free on Devin paid tiers through early October, then $3/$15 list pricing. No public API or model card β SWE-2 is a product model, not a commodity. The open frontier keeps closing: the base model alone was already #7 overall.
GPT-6 Astra β OpenAI's $10/$50 flagship
OpenAI announced GPT-6 Astra on Sep 3 and shipped the API model Sep 4. $10/$50 per 1M β 2.5x GPT-5.6 Sol's $4/$20 promo β with $1 cached input, $12.50 cache writes, and long-context rates above 272K tokens (2x input/cache, 1.5x output). Batch/Flex at 50%, Fast mode at 2x. 1.05M context, 128K output. Saturates ARC-AGI-3 (99.9%) and FrontierMath Tier 4 (98%); ExploitBench 100% without safeguards. First model to hit OpenAI's critical cybersecurity threshold. AA Intelligence Index: 61, but ~70% more token-efficient than Sol β less than half Fable 5's cost per coding task.
Fable 5.1 β the cache-read cut is the story
Claude Fable 5.1 (Sep 1) keeps $10/$50 but cuts cache reads 75% to $0.25/1M. Anthropic's own usage data: ~25% total savings on typical workloads, up to ~45% for agentic work where cache hits dominate. Terminal-Bench-Science 52.6% (more than doubles Fable 5), AutomationBench 31.4% (vs 17.1%), AA Index 66 β #1. Candid caveats from the system card: forced tool use now returns 400, thinking blocks are model-bound, and Mythos 5.1 is 'less honest under pressure' than recent Claude models.
Gemini 3.8 Flash β fourth Flash in four months
Google shipped Gemini 3.8 Flash on Sep 2 β the fourth Flash since May (3.5, 3.6, 3.7, 3.8). AA Index 59 (+3 over 3.7) at the same $0.75/$3.75 intro through Dec 31, then $1.50/$7.50. ~300 tok/s β the fastest output speed Artificial Analysis has measured. Trained on long-running agentic loops, it 'works harder': ~40% higher cost per task than 3.7 (30% more output tokens), so use low effort for efficiency-first routing. Gemini 3.8 Flash Cyber reaches trusted defenders via the Fairwind Program β 2.6x more correct Chrome patches than larger frontier models.
Muse Spark 1.3 β cheapest model at the frontier
Meta's Muse Spark 1.3 (Sep 2) scores 61 on the AA Intelligence Index at xhigh (62 at max, partner preview) β behind only Fable 5.1 (66) and Opus 5 (63). $1.25/$4.25 unchanged, $0.55 per Index task vs $0.94β0.95 for GPT-5.6 Sol and Grok 4.6 β cheapest of any model at 59+. #1 on Tau3-Bench Banking. Meta teases open weights and larger Muse models around Q4.
The price floor keeps collapsing
Three weeks after four frontier launches (Sep 1β3), the commercial floor dropped again: DeepSeek V4.1-Flash (Sep 10) at $0.15/$0.60 off-peak with cache hits at $0.003/1M β 50x below its own cache-miss rate β and Inception's Mercury 2.5 diffusion LLM (Sep 8) at $0.04/$0.15 launch pricing running 1,107 tok/s. Cognition's SWE-2 (Sep 10) matches Fable 5.1 within a point on FrontierCode at ~64% lower cost. GLM-5.3-Flash returned to list ($0.15/$0.50) on Sep 10. Calendar: GPT-5.6 Sol promo ends Nov 21 ($4/$20 β $5/$30), Gemini 3.8/3.7 intro ends Dec 31 (doubles Jan 1, 2027), Mercury 2.5's 80% discount ends early Oct.
Grok 4.6 β frontier at $2/$6
xAI's Grok 4.6 (Aug 12) launched with Cursor at $2/$6 per 1M, $0.50 cached input, with a 1753 ELO claim that overtakes Kimi K3. Live on xAI API, Grok Build, Cursor, and Grok Bot, plus partners OpenRouter, Vercel, and Cloudflare. A fast variant costs twice the price. xAI frames it as roughly half the cost of other frontier models. 2x included usage in Cursor and Grok Build for the first week.
GLM-5.3 β 743B open-weight coder, weights now public
Z.ai (formerly Zhipu AI) released GLM-5.3 on Aug 14 β a 743-billion-parameter model built by scaling post-training on the GLM-5.2 base. Weights went public on Hugging Face Aug 27 with 51 community quantizations within a week. AA Intelligence Index v4.3 (max): 44.9 β the strongest open-weight model on the index, ahead of Kimi K3 (43.8). GLM-5.3-Flash ($0.15/$0.50, 1.31M ctx) measures 41.9 β the cheapest per-token open API. 'Dramatic improvement over GLM-5.2 with fewer output tokens.'