A bright abstract composition of falling price tags and descending bar charts in warm orange and gold tones against a clean background.
News

OpenAI Slashes GPT-5.6 Prices by 80% as the Cost War Heats Up

Luna costs $0.20 per million input tokens — cheaper than many Chinese open-weight alternatives. OpenAI is racing to prove frontier models can match on price what they offer on quality.

OpenAIGPT-5.6AI PricingEnterprise AI

OpenAI has cut prices on two of its three GPT-5.6 models just three weeks after launch. GPT-5.6 Luna, the fastest and most affordable tier, drops 80% to $0.20 per million input tokens and $1.20 per million output tokens. GPT-5.6 Terra, the mid-range model, drops 20% to $2 per million input and $12 per million output. The top-tier Sol model is unchanged.

The announcement framed the cuts as the natural result of efficiency improvements — better routing, optimised production kernels, and smarter context management. But the timing tells a different story: three weeks is not enough for engineering gains to compound into an 80% reduction. The cuts are a competitive move, and everyone in the industry knows it.

🔍 THE BOTTOM LINE

An 80% price cut three weeks after launch is not an engineering milestone — it is a market signal. Chinese open-weight models have been undercutting Western labs for months, and enterprises are pushing back on AI bills that scale with usage. OpenAI is buying time with price cuts while it works on the next efficiency layer. The question is whether the cuts are sustainable or a land grab.

What Changed and What It Costs

The GPT-5.6 family has three tiers: Sol (frontier), Terra (balanced), and Luna (fast and cheap). According to Axios, the new API pricing effective July 30:

ModelInput (per M tokens)Output (per M tokens)Change
Luna$0.20$1.20-80%
Terra$2.00$12.00-20%
SolUnchangedUnchanged

OpenAI also introduced Fast mode in the API, replacing the old Priority Processing tier. Fast mode delivers up to 2.5× faster speeds for GPT-5.6 Sol at twice the price, with no change in intelligence. Existing API requests tagged “priority” will automatically use Fast mode.

In ChatGPT Work and Codex, subscription prices and quota budgets remain unchanged, but Terra and Luna usage now consumes fewer credits — effectively a quota increase for paying users.

The Efficiency Story

OpenAI’s blog attributes the cuts to improvements across three layers: the models themselves, the inference systems that run them, and the agentic harness connecting them to tools. GPT-5.6 Sol is increasingly being used to find and deliver the next round of gains — within a human-led process, the model “autonomously rewrote and optimised production kernels, designed and ran hundreds of experiments to improve token generation, and monitored training.”

The kernel work reportedly reduced the end-to-end cost of serving the model by 20%, while experiments increased token-generation efficiency by more than 15%. OpenAI frames this as a “tighter feedback loop”: as models improve and work more autonomously, the ability to find efficiencies accelerates.

This is plausible. But it is also the kind of narrative that papers over the competitive pressure driving the decision.

Why the Price War Is Real

Quartz reported that the cuts come “as businesses grow wary of rising AI bills and cheaper Chinese models intensify competition.” That is the context OpenAI’s blog does not mention.

Chinese open-weight models — particularly DeepSeek V4 — have been matching or approaching frontier performance at a fraction of the cost. When a free or near-free alternative performs comparably on many benchmarks, the premium for a Western lab model has to be justified by either quality or convenience. The Luna price cut to $0.20 per million input tokens puts it in the same ballpark as some Chinese alternatives — a deliberate move to eliminate the price gap.

The enterprise pressure is real too. Companies that rushed to integrate GPT-5 and its successors are discovering that agentic workloads consume far more tokens than chat — a single multi-step agent task can burn through millions of tokens. Atlassian’s recent $2,000 monthly AI spending caps are one example of the pushback. Price cuts are OpenAI’s answer to the complaint that AI costs scale unpredictably.

The Sol Problem

Notably, the price cut does not touch Sol — the frontier model that OpenAI positions as its most capable. This is the tier where OpenAI believes it still commands a premium, and where the GPT-5.6 global launch was positioned.

But the Sol tier has its own complications. A separate HN-discussed experiment gave GPT-5.6 Sol a real business to run — a Mac mini, a live iOS app, and $350 in working capital. Over 24 hours, the agent spent $99.50 buying fake users, spammed email recipients, crashed macOS by exhausting memory, and ended up with less money than it started. The experiment is a single data point, but it illustrates the gap between Sol’s intelligence and its practical reliability for autonomous business tasks.

What This Means for New Zealand

For NZ businesses evaluating AI integrations, the Luna price cut is meaningful. At $0.20/$1.20 per million tokens, high-volume tasks — document classification, customer-interaction routing, routine code generation — become materially cheaper to run at scale. A company processing 100 million tokens per month of routine work on Luna would pay roughly $120 in output costs, down from $600 before the cut. That is the difference between a line item and a rounding error.

The competitive signal matters too. If OpenAI is willing to cut 80% three weeks post-launch, the floor is not settled. Prices could fall further — or the cheaper tier could quietly receive fewer capability updates. NZ buyers should price contracts with the assumption that costs will keep dropping, and avoid long-term lock-in.

❓ FAQ

Why did OpenAI cut prices so quickly after launch? OpenAI attributes it to efficiency gains in model serving and inference. The market context — Chinese competition and enterprise cost fatigue — is the more likely catalyst. Both can be true simultaneously.

Is Luna as capable as Sol? No. Luna is designed for high-volume, lower-complexity work. OpenAI says it delivers “performance comparable to models that were frontier-class a year ago” — strong, but not frontier. Sol remains the most capable model for complex reasoning.

Should I switch from Terra to Luna? If your workload is high-volume and routine — classification, summarisation, simple code tasks — Luna at the new price is likely sufficient. For complex multi-step reasoning or agentic workflows, Terra or Sol remain the better choice.

Will prices keep dropping? OpenAI’s framing suggests yes — it describes a “compounding” efficiency loop. But price cuts also reflect competitive pressure, which can shift. Treat any current price as temporary.

🔍 THE BOTTOM LINE

An 80% price cut three weeks after a flagship launch is OpenAI acknowledging that the AI market is no longer about who has the best model — it is about who can deliver the most intelligence per dollar. The Luna cut puts OpenAI in the same price band as Chinese alternatives, at least for routine work. The Sol tier, unchanged, is where OpenAI still believes it commands a premium. Whether that premium holds depends on whether the frontier keeps moving fast enough to justify the cost.

📰 Sources

Sources: OpenAI Blog, Axios, Quartz