There’s now an index for it, and on Monday the index hit rock bottom. The LLM Token Expenditure Index, which tracks the going market rate for a large-language-model token, fell to 97 US cents on Monday 1 September — its lowest reading since the index was created late last year, and less than half its high from earlier this summer, CNBC reported.
What the number actually measures
Most people never see token prices directly; they’re buried in API bills. Silicon Data’s index aggregates the going market rate for LLM tokens across providers into one daily figure — think of it as a commodity price for intelligence. When it falls, running a chatbot query gets cheaper for everyone downstream.
Three forces are pushing it down, according to Charles-Henry Monchau, chief investment officer at Syz Group, who wrote about the drop on Tuesday:
- Chinese open-weight models. Moonshot’s Kimi K3 and its peers can undercut frontier-lab prices, and buyers are switching. The capability gap between open Chinese models and the Western frontier is now measured in months, not years.
- Frontier price cuts. OpenAI cut prices on two GPT-5.6 models in late July — an 80% cut in the case we covered at the time. DeepSeek went further, locking in a permanent 75% discount on its flagship. When the price leader cuts, everyone follows.
- Dynamic pricing. More labs are letting rates rise and fall with demand, airline-style. That smooths revenue but pushes the market-clearing price down.
The squeeze nobody priced in
Here’s the uncomfortable part for the model labs: token deflation “compresses the revenue line while compute commitments stay fixed,” as Monchau put it. Labs have signed up for billions in GPU capacity at fixed cost. If the price per token keeps falling faster than usage grows, the margin math stops working — right as OpenAI and Anthropic prepare for public listings after confidentially filing with regulators this summer.
Silicon Data’s head of research, Steve Hou, offered the blunter interpretation: prices this low may signal there’s already enough supply “to provide sufficient capabilities for most tasks.” In a market with a genuine frontier premium, prices wouldn’t collapse like this. The moat, Monchau argues, has to shift from raw capability toward distribution, memory and context — the things you can’t copy by downloading weights.
We’ve been tracking this all year: the permanent DeepSeek discount, GLM running at half the cost of Blackwell inference, corporate budgets blowing up on trivial token burn. The pattern keeps pointing the same way. Intelligence is behaving like a commodity, and commodity markets are cruel to sellers.
What it means for everyone else
For users, this is unambiguously good — the price of running AI workloads keeps falling, and small companies building on these APIs are the beneficiaries. For the labs, it’s a vise: their costs are largely fixed, their revenue per unit is falling, and their response has been to race toward scale and lock-in before the price of intelligence reaches whatever floor the market decides on.
Hedging appropriately: none of this means AI companies are about to collapse. Demand is enormous and growing. But it does mean the easy story — “charge premium prices for the smartest model” — is dying, and the companies priced for endless margin expansion may need to explain something harder to investors: what exactly justifies the multiple when the product’s price keeps falling?
For New Zealand businesses, this is the good kind of deflation. Whatever you’re paying for AI today — API calls, coding assistants, customer service bots — expect that number to keep falling. The companies getting hammered by it are the ones with billions in fixed compute. If you’re buying rather than selling, this market works in your favour.
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.