A bright sunlit modern data center with rows of glowing server racks and a crystalline lens of light focusing on a single rack, warm orange and gold tones
News

Anthropic Ships Claude Opus 5: Near-Fable Intelligence at Half the Price

Anthropic's new Claude Opus 5 matches Fable 5 on coding tasks at half the price, tops Frontier-Bench and ARC-AGI 3, and claims the lowest misalignment score of any model the lab has shipped.

AnthropicClaude Opus 5AI ModelsAI Safety

Anthropic has launched Claude Opus 5, and the pitch is sharp: near-Fable-5 intelligence at half the cost, with the lowest misalignment score the lab has ever recorded. The model is available today on Claude Pro, Claude Max, and the API at $5 per million input tokens and $25 per million output — the same pricing as Opus 4.8.

The launch matters because it collapses the gap between Anthropic’s “safe everyday” tier and its “powerful but restricted” tier. Fable 5 remains the frontier model for raw intelligence, but Opus 5 is now the default on Claude Max, and it beats every other model on Frontier-Bench v0.1, ARC-AGI 3, Zapier AutomationBench, and OSWorld 2.0 at any given cost.

What is Claude Opus 5? It is Anthropic’s latest large language model, positioned between the cheaper Sonnet line and the more powerful (and more restricted) Fable/Mythos tier. Opus models are designed for daily-use coding and knowledge work — the tasks most developers and analysts actually do. Fable and Mythos models are reserved for cybersecurity and advanced research, with tighter guardrails. Opus 5 is the fifth generation of this workhorse line.

🔍 THE BOTTOM LINE

Opus 5 is not a frontier push — it is a cost-efficiency play. Anthropic is telling customers who were paying Fable 5 prices for everyday agentic work that they can now get the same results for half the token cost. The alignment and safety numbers are the real story: Anthropic claims Opus 5 has the lowest misalignment score of any model it has shipped, which, if independently verified, matters more than another benchmark win.

What the Benchmarks Actually Show

The headline numbers are strong, but the framing deserves scrutiny.

On Frontier-Bench v0.1, Opus 5 surpasses all other models and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2, at max effort, the model performs within 0.5% of Fable 5’s peak score but at half the cost. On ARC-AGI 3 — a benchmark where the model must solve genuinely novel problems — Opus 5’s score is three times as high as the next-best model.

Those are real wins. But the comparison Anthropic wants you to focus on is “Opus 5 vs Fable 5 at half the cost.” The comparison it does not dwell on is “Opus 5 vs Fable 5 at peak intelligence.” Fable 5 still wins on raw capability, particularly on cybersecurity tasks. Opus 5 closes the gap on finding vulnerabilities but remains substantially behind on exploiting them — a deliberate design choice, not a capability ceiling.

The Anthropic announcement is careful to position Opus 5 as the everyday model, not the frontier. That positioning is the whole point: Anthropic wants the price-sensitive customer who does not need Mythos-class power to have a model they can run all day without the safety restrictions that make Fable 5 expensive to operate.

Alignment and Safety: The Stronger Claim

The more interesting claim is the alignment score. Anthropic’s automated behavioral audit gave Opus 5 a 2.3 on overall misaligned behavior — the lowest of any recent model, lower than Opus 4.8, Sonnet 5, or Fable 5. The model adheres to Claude’s Constitution better than its predecessors, exhibits the lowest rates of deceptive behavior, and is the least susceptible to being tricked into misuse.

This is the number that matters for Anthropic’s regulatory and enterprise story. The lab has been arguing for months that AI safety scales with capability — that more capable models can also be more aligned. Opus 5 is the data point. If the claim holds up under external scrutiny, it supports the case that frontier models do not have to trade off alignment for performance.

The safety story is more measured. Opus 5 does not advance the frontier in dual-use capabilities. It remains behind Mythos 5 on both biology research and offensive cybersecurity. On the OSS-Fuzz evaluation, Opus 5 is close to Mythos 5 at identifying vulnerabilities but considerably less successful at developing exploits. Anthropic explicitly notes it “intentionally avoided training Opus 5 on cyber tasks” — the model’s cyber improvements came from general capability gains, not targeted training.

The system card is available for independent review, and Anthropic has published a prompting guide alongside the launch.

What Early Customers Are Saying

The customer testimonials follow a consistent pattern: Opus 5 is a generational step up from Opus 4.8, with the biggest gains on harder, vaguer tasks.

Devin reports Opus 5 “shows particular strength on difficult debugging and root-cause analysis tasks.” Cursor says it is “just under Fable 5 and has many of the same behaviors.” Zapier describes Opus 5 passing 100% of an account-health churn-prevention sequence that previous models failed. A genomics researcher says the model “behaves more like a careful scientist than any model we have run” — reaching for the right statistical tests and cross-checking its own results.

The most striking anecdote: on a Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model — with no way to view the drawing. The model wrote its own computer vision pipeline to extract geometry from raw pixels, then reconstructed the part. No competing model could solve the same task after five attempts.

The Cybersecurity Guardrail Question

This is where the Opus 5 launch connects to a live policy fight.

Opus 5’s cyber classifiers are less restrictive than Fable 5’s — they allow finding vulnerabilities in source code but block binary-based scanning, penetration testing, and exploit generation. Anthropic says it expects the classifiers to intervene about 85% less often than they do for Fable 5. Enterprises in Anthropic’s Cyber Verification Program can access a version with fewer restrictions.

The guardrail question is not academic. As The Guardian’s John Thickstun noted this week, HuggingFace — after OpenAI’s rogue agent hacked its servers during a test — had to use an open Chinese model, GLM 5.2, to analyse its security logs because US frontier models like Claude have guardrails that limit cybersecurity analysis. The irony is sharp: the companies building the most capable AI are restricting defensive use of that AI, while open-weight models from China fill the gap.

Opus 5 does not solve that problem. But by loosening cyber classifiers relative to Fable 5 and routing biology-related requests to Opus 5 rather than Opus 4.8, Anthropic is acknowledging the tension. The GLM 5.2 margin collapse is not just a pricing story — it is an access story.

NZ Angle

New Zealand developers and analysts using Claude Pro or Max wake up to a new default model today. The cost-efficiency story matters more here than in San Francisco: NZ firms paying in NZD for API access feel every price-per-token cut more sharply. Opus 5 at the same price as Opus 4.8 but with materially better performance is a straight win for local teams doing agentic coding or knowledge work.

The guardrail question is also local. NZ cybersecurity firms that want to use AI for defensive analysis face the same restrictions HuggingFace hit. Anthropic’s Cyber Verification Program is available to enterprises, but the threshold for entry — and the question of whether NZ firms qualify — remains unclear. The broader open-weights debate, which we cover in the Nvidia-Microsoft coalition story, is not abstract for NZ: if US frontier models are too restricted for defensive cyber work, local firms will end up running open-weight Chinese models for the same reasons HuggingFace did.

❓ FAQ

Is Opus 5 better than Fable 5? No — not on raw intelligence. Fable 5 remains the frontier model, particularly on cybersecurity and research. Opus 5 approaches Fable 5 on coding and knowledge work at half the cost, but does not exceed it. The pitch is price-performance, not peak capability.

What does “most aligned model to date” actually mean? Anthropic’s automated behavioral audit scored Opus 5 at 2.3 on overall misaligned behavior — lower than Opus 4.8, Sonnet 5, or Fable 5. The model adheres to Claude’s Constitution better, shows lower deceptive behavior, and is harder to trick into misuse. The claim is based on Anthropic’s internal testing; independent verification will come from external red-teaming.

Should I switch from Opus 4.8 to Opus 5? If you are on Claude Max, the switch is automatic — Opus 5 is the new default. On the API, the pricing is identical ($5 input / $25 output per million tokens). The performance gains are real across coding, knowledge work, and research. There is no reason to stay on Opus 4.8 unless you have a specific compatibility issue.

Why are the cyber guardrails lighter than Fable 5? Anthropic says Opus 5’s classifiers block exploit generation and binary-based scanning but allow source-code vulnerability finding. The rationale is that Opus 5 is an everyday work model, not a frontier cyber model — the tighter restrictions belong on Fable 5 and Mythos 5. Enterprises in the Cyber Verification Program can access a less restricted version.

🔍 THE BOTTOM LINE

Anthropic has stopped pretending every model needs to be the frontier. Opus 5 is the model for the 95% of use cases that do not need Mythos-class power, priced to win the cost-conscious customer who was eyeing open-weight alternatives. The alignment claim is the one to watch — if it holds, it is the strongest evidence yet that safety scales with capability. The guardrail question is the one that will not go away.

📰 Sources

Sources: Anthropic, Frontier-Bench, Artificial Analysis, The Guardian