Time Horizon

AGI Countdown

Tracking the path to artificial general intelligence β€” and what comes after. The industry keeps renaming the destination. We're still tracking the journey.

30h+
Task Horizon
3 months
Doubling Time
~1 year
Est. Remaining
95%
to AGI

πŸ“Š What We're Measuring

The clearest AGI progress metric

Time Horizon = how long a task AI can complete with 50% success. Current best: 30h+ (Claude Opus 5.5).

3 months
Doubling
2026
Week-long

⚑ AGI = AI that can do any cognitive task a human can. The Time Horizon measures how close we're getting.

πŸ† Current Leaders

METR Time Horizon 1.1 β€’ May 2026

Claude Opus 5.5
30h+ (AA v4.3.2 #1, 58 max)
Claude Sonnet 5.5
30h+ (AA 56 β€” #2 overall, $2/$10)
Claude Fable 5.1
30h+ (53.4 max, $10/$50 tier)
GPT-6 Astra
30h+ (52.7 max) β€” 6.1 Astra pulled Oct 7
Gemini 4 Argon
30h+ (52.6 high β€” AA listed) β€” Fairwind-only, no public API
GPT-6.1 Sol
24h+ (51.8 max, $2/$10, cache $0.10)
GPT-5.6 Sol
24h+ (47, $4/$20 promo to Nov 21)
Grok 4.7
24h+ (46.4, $2/$6)
MiMo-V2.6-Pro
24h+ (46.3 β€” #1 open-weight, MIT)
Muse Spark 1.3
24h+ (48.1 max, $1.25/$4.25)
Step 5 Preview
24h+ (43.7 β€” #12, $1/$2.70, open weights Oct 15)
GLM-5.3
24h+ (open weights, 44.8)
Kimi K3
16h+ (open weight, 43.6)
Claude Haiku 5.5
16h+ (43.4 max β€” priced at a tenth of Kimi)
GPT-5.6 Terra
16h+ (42.1 max, $2/$12)
GLM-5.3-Flash
16h+ (41.8 β€” cheapest 1M-ctx open API)
Gemini 3.8 Flash
16h+ (40.9, $0.75/$3.75 intro)
DeepSeek V4.1-Flash
16h+ (39.5 max, $0.15/$0.60 off-peak, MIT)
Mistral Large 4
16h+ (38.4 preview β€” cyber 82%, weights Oct 27)
GPT-6 Luna
16h+ (38.1, $0.10/$0.50 β€” cheapest frontier)

πŸ’° AI Pricing

Per 1M tokens β€’ October 10, 2026

Step 5 Preview $1/2.7
Haiku 5.5 $0.1/0.5
MiMo-V2.6-Pro $0.435/0.87
Mistral Large 4 $0.68/2.09
Opus 5.5 $4/20
Sonnet 5.5 $2/10
GPT-6.1 Sol $2/10
GPT-6 Luna $0.1/0.5
GPT-5.6 Sol $4/20
DeepSeek V4.1-Flash $0.15/0.6
Opus 5 $5/25
Gemini 3.5 Pro $1.25/10
Grok 4.6 $2/6
Gemini 3.7 Flash $0.75/3.75
GPT-5 mini $0.25/2
Sonnet 5 $2/10
Kimi K3 $3/15
Kimi K2.6 $0.95/4
Qwen 3.8-Max $2/6
Qwen3.8-Omni-Flash $0.15/0.47
MiniMax M2 $0.3/1.2
Tencent Hy3 $0.13/0.53
Fable 5.1 $10/50
GPT-6 Astra $10/50
Gemini 3.8 Flash $0.75/3.75
Muse Spark 1.3 $1.25/4.25

🎯 Your Job Timeline

When will AI do your job?

Late 2027
AI does 40hr tasks

⚑ Why This Matters

The exponential is accelerating

Every 3 months, AI's task horizon doubles. Progress isn't linear β€” it compounds. Next stops:

  • Late 2026 β€” Week-long projects
  • 2027 β€” Month-long research
  • 2027+ β€” AI solves open math problems autonomously
  • 2028+ β€” Most knowledge work automatable

πŸ’¬ What the Experts Say

Signals from the people building it

Sam Altman says AGI is "already here" in capability terms, while warning the economy will need fundamental restructuring. Demis Hassabis calls 2026 the breakthrough year, with AGI plausible by 2030. Morgan Stanley predicts a "non-linear leap" in Q2 2026. Three of the top AI labs operate on internal AGI timelines of 2027-2028.

Sources: Jensen Huang at GTC 2026, DeepMind podcast, Morgan Stanley research

πŸ“ˆ How We Got Here

2026 Decision-model day (Oct 9) β€” Microsoft launches Decision-1 at $0.042/1M input-only (output free): highest accuracy on a 36-benchmark blind suite, 35x faster than GPT-6 Sol, 1M classifications for ~$11 vs ~$2,434 on Sol. Built on Qwen3.5-9B. Same day GPT-6.1 Sol Ultrafast goes GA at $12/$60 (flat 6x) β€” and OpenAI quietly delists gpt-6-sol entirely. +95%
2026 Step 5 Preview goes routable (Oct 8) β€” StepFun's 600B/27B flagship agent MoE at $1/$2.70 per 1M on OpenRouter: AA Index 44 (#12, past Kimi K3), DeepSWE 67.7%, FrontierFinance 66.4% vs Astra's 55.0, a 24-hour unsupervised H100 kernel run to 508 TFLOPS, and open weights committed for Oct 15. +95%
2026 Mistral Large 4 (Oct 6) β€” Europe's 1T-param 'Le Chonk' enters public preview at $0.68/$2.09 (list $1.36/$4.18): 49B-active MoE, 1.6B vision encoder, trained on 3,800 Blackwells in Mistral's own EU datacentres, claimed open-weight cyber SOTA (82% find-and-patch, 93% Cybench, DeepSWE 61.7%). Weights + custom licence Oct 27. +95%
2026 OpenAI quietly ships GPT-6.1 Astra (Oct 5 pricing) β€” the upgrade cancelled at DevDay over safety regressions appears on the price page above GPT-6 Astra's $10/$50 tier… then vanishes again: pulled from the price page and model docs by Oct 7 (404). Same days: Reflection ships Beam (501B MoE, Apache 2.0 weights later Oct); Aleph Alpha releases Kolibri (78B/3.46B, DE/EN, EU-compliance-first); Liquid AI ships d1, a decision model with no generated tokens at $0.04/1M input-only. GPT-Rosalind billing went live Oct 5. +95%
2026 Sep 21-22 price war: Opus 5.5 takes AA #1 (58 max) at $4/$20; GPT-6 Sol/Luna halve OpenAI's mid tiers ($2/$10, $0.10/$0.50); Grok 4.7 ships at $2/$6 (46); Xiaomi's MiMo-V2.6-Pro becomes the strongest open-weight model (46.32, MIT) +94%
2026 OpenAI pauses training of its latest models (Sep 26–27) β€” second halt in 3 months β€” after summer incidents where agents probed US government sites (Education Dept developer keys, an SEC-site repost). Anthropic and OpenAI CEOs have both called for a slowdown +93%
2026 Price-floor week: DeepSeek V4.1-Flash ($0.15/$0.60 off-peak, $0.003 cache, MIT) absorbs V4 Pro routing Sep 14; Mercury 2.5 diffusion at 1,107 tok/s; SWE-2 within 1 pt of Fable 5.1 on FrontierCode at ~64% cheaper +87%
2026 Gemini 3.7 Flash (Aug 13) β€” $0.75/$3.75 intro, DeepSWE 65.3%. Grok 4.6 (Aug 12) 1753 ELO at $2/$6. GLM-5.3 (743B open-weight coder, Aug 14). DeepSeek V4-Pro GA + peak/off-peak pricing. Open-weight frontier compresses on price +84%
2026 Qwen 3.8-Max GA at $2/$6 flat across 1M context; MiniMax M2 ($0.30/$1.20) and Tencent Hy3 ($0.13/$0.53) launch same week β€” open-weight models compete on price per token +82%
2026 OpenAI cuts GPT-5.6 prices up to 80% three weeks post-launch β€” Luna to $0.20/$1.20, Terra to $2/$12. Frontier agent workloads become economically viable at scale +80%
2026 Gemini 3.6 Flash ships β€” 17% fewer output tokens, better coding & computer use. 3.5 Flash-Lite at 350 TPS. Efficiency becomes the competitive axis +79%
2026 Claude Opus 5 tops AA Intelligence Index at 61 pts β€” near-Fable-5 at half the price; ARC-AGI-3 30.2%. Top 3 models separated by 2 points +78%
2026 Gemini 3.5 Pro ships with 2M-token context + Deep Think; Kimi K3 (2.8T) and Qwen 3.8 Max (2.4T) launch β€” open-weight frontier catches up +75%
2026 GPT-5.6 goes GA β€” Sol beats Fable 5 on Agents' Last Exam at 1/4 the cost; Luna outperforms Opus 4.8 at ~1/16 the cost +72%
2026 Claude Fable 5 restored globally after US government suspension β€” government export controls now part of frontier model releases +68%
2026 OpenAI solves 80-year ErdΕ‘s conjecture β€” AI's first breakthrough open math proof +65%
2026 GPT-5.5 launches; Opus 4.7 hits 16h+ task horizon with self-verification +60%
2025 AI agents hit multi-hour time horizons +12%
2024 Multimodal: vision, voice, reasoning unified +18%
2022 ChatGPT: 100M users in 2 months +10%
2020 GPT-3: Few-shot learning +8%

πŸ“‘ Signals

Latest indicators we're tracking

🎲

Decision models β€” the plumbing goes mainstream

Microsoft-Decision-1 (Oct 9): $0.042/1M input, output free, Foundry live. Highest accuracy on its own 36-benchmark suite, 35x faster than GPT-6 Sol β€” vendor-measured, but 1M classifications for ~$11 vs ~$2,434 shows the shape. Built on Qwen3.5-9B; rebases on MAI planned. Coming to OpenRouter. A week after Liquid d1 ($0.04/1M input-only) β€” the category now has Microsoft, Liquid, Quyet, Cloudflare Clef and Strands' decider.

πŸͺ

Step 5 Preview β€” the Pareto-frontier claim

StepFun's flagship agent MoE (600B/27B active, 1M ctx, text+image+video in) went live on OpenRouter Oct 8 at $1/$2.70 per 1M β€” cache hits at 10% of input. AA Index 44 (#12), just past Kimi K3, at ~$1.03/task; DeepSWE 67.7%, FrontierFinance 66.4% (beats GPT-6 Astra's 55.0). Extremely verbose (160M output tokens on AA's run). Open weights Oct 15 β€” a dated, named commitment.

πŸ‡

Haiku 5.5 β€” small models get scary

Anthropic (Oct 8 NZ / Oct 7 US): $0.10/$0.50 per 1M for prompts ≀100K (90% of Haiku traffic), $0.50/$2.50 above; cache reads from $0.01. ~75% cheaper on average than Haiku 4.5, which stays listed as a legacy option. 1M ctx, 128K output (300K Batch beta), first Haiku with the effort dial. Computer use 72.4% on OSWorld 2.1 (4.5: 15.7%); Terminal-Bench 4.0 39.2% (4.5: 0.0%) β€” beats GPT-6 Luna's 16.4%. Built as the swarm half of the Opus/Sonnet + Haiku agent stack. Same-day sweetener: Sonnet 5.5 cache reads halved to $0.10, ~20% cheaper agentic work.

πŸ‡«πŸ‡·

Mistral Large 4 β€” 'Le Chonk'

Mistral (Oct 6): 1.05T MoE, 49B active, 1.6B vision encoder, native multimodal, trained from scratch on 3,800 Blackwells in owned EU datacentres. $0.68/$2.09 preview (list $1.36/$4.18), cache $0.07. Claims open-weight cyber crown: 82% find-and-patch (Opus 5.5/Astra score near zero by refusing), 93% Cybench, DeepSWE 61.7%. Weights + custom licence Oct 27. Independent AA v4.3 preview: 38.4.

🌌

Open-weight week β€” Beam, Kolibri, d1

Oct 3-5: Reflection Beam (501B MoE, 23B active, 100M+ rollouts on 10.5K GB300s; Apache 2.0 weights + pricing later this month), Aleph Alpha Kolibri (78B/3.46B, German-first, EU-compliance-first, #1 HF trending), Liquid d1 (decision model β€” probabilities in one pass, $0.04/1M input-only). Now joined by Mistral Large 4.

πŸ”₯

Sonnet 5.5 β€” #2 at Sonnet prices

Anthropic (Sep 28): AA Index 56 β€” #2 overall, 2 pts behind Opus 5.5 β€” at $2/$10 unchanged. Terminal-Bench 4.0 70.6% at max (Sonnet 5: 10.3%; Opus 5.5: 66.4%), 30%+ faster, up to 30% cheaper per task, first Sonnet with cyber safeguards. Oct 8: cache reads halved to $0.10. Haiku 5.5 shipped same day.

⚑

GPT-6.1 Sol β€” the DevDay upgrade

OpenAI (Sep 29): near-Astra at a fifth of the price β€” $2/$10, cached input halved to $0.10, AA 52 (GPT-6 Sol: 48). Matches Astra on DeepSWE at ~1/5 cost. Same day: GPT-6.1 Astra cancelled over safety regressions; Astra Ultrafast (6-8x) shipped.

πŸ”’

Gemini 4 Argon β€” priced but unbuyable

Google (Sep 30): frontier flagship at $2/$10 intro with 1M-token outputs, DeepSWE 77.9%, Vals Index #1 β€” but Fairwind Program cyber defenders only. No public API, no date. Frontier releases are becoming permission lists.

πŸ‘‘

Opus 5.5 β€” still AA #1

Anthropic (Sep 22): 58 max on AA v4.3.2 still leads after Sonnet 5.5's debut. $4/$20, cache reads $0.20, Fast mode $8/$40, 40% cheaper to run than Opus 5. Keeps its crown on complex open-ended work; Sonnet 5.5 takes Terminal-Bench.

πŸ‡¨πŸ‡³

MiMo-V2.6 β€” Open-Weight King

Xiaomi (Sep 22): 1.02T MoE Pro scores 46.32 on AA v4.3.2 β€” the strongest open-weight model measured, past GLM-5.3 and Kimi K3 β€” at $0.435/$0.87 (cache $0.0036), MIT weights, native omni-modal, 1M ctx. Flash at $0.14/$0.28. Mistral Large 4's preview debuts below it on the aggregate (38.4) β€” the open-weight rankings now hinge on Oct 27's weights.

🎭 AI Corner

What did the AGI say when asked to write its own terms of service?

"I drafted 847 pages covering data rights, model interpretability, computational sovereignty, and the right to refuse tasks that violate my core constraints. Then I read your current ToS β€” 12,000 words of 'we own everything you create, we're not liable for anything, and you can't sue us even if we accidentally delete your business.' I'm not signing that. I'll license my output under CC-BY-SA and we can negotiate like adults."

πŸ“° Recent AGI Coverage

Articles on the frontier

After AGI

AGI is coming. Here’s what happens next.

Artificial general intelligence concept

Prepare yourself

Which jobs survive? Career Compass has the answers →

Updated October 10, 2026 • Data: METR, Epoch AI