A glowing silicon microchip on a bright circuit board, warm golden light, circuit traces flowing into neural network patterns, vibrant orange and teal tones.
News

AMD Bought the Chip That Turns an AI Model Into Silicon — and the Memory Market Should Read Some History

AMD just bought Taalas, the startup that etches entire AI models into silicon. The HC1 runs Llama 3.1 8B at up to 17,000 tokens per second without touching HBM. It sounds like nothing — until you remember what smartphones did to cameras and Uber did to taxi medallions.

AMDTaalasAI ChipsMemory MarketInference

On August 6, 2026, AMD announced it had acquired Taalas, a 24-person Toronto startup that does something no GPU does: it etches an entire AI model directly into silicon. No memory fetches. No HBM stacks. The weights of Meta’s Llama 3.1 8B are burned into the transistors themselves, and the chip answers at up to 17,000 tokens per second — company figures, but enough that AMD decided ownership beat competition.

The acquisition barely made headlines outside the chip press. It should have, because it points at a question every booming market eventually has to answer: what happens when a technology arrives that your entire business model assumes is impossible?

🔍 THE BOTTOM LINE

The global memory market is enjoying the biggest boom in its history because AI inference is believed to need staggering quantities of DRAM and HBM forever. Hardwired inference chips remove that assumption at the architectural level — the model never leaves the chip, so it never queues for memory. This does not mean a RAM crash is coming tomorrow. It means the “impossible” scenario has a mechanism, a price tag, and now a giant corporate owner. Markets have a habit of pricing the status quo as permanent right up until it isn’t.

What AMD Actually Bought

Taalas was founded in 2023 by Ljubisa Bajic, a former AMD and Nvidia architect who co-founded Tenstorrent and served as its first CEO before leaving in early 2023. In February 2026 the company unveiled its first chip, the HC1, and raised $169 million to build it out. According to Bajic, the whole thing was developed by 24 people on about $30 million of engineering spend — pocket change by AI chip standards.

The HC1’s numbers, per the company’s own claims reported by Data Center Dynamics and Forbes: built on TSMC’s 6nm process, 815 square millimetres, 53 billion transistors, running Llama 3.1 8B at up to 17,000 tokens per second while drawing roughly a tenth of the power of an Nvidia B200 doing the same job, as The Register reported at the time of the acquisition. No high-bandwidth memory. No liquid cooling. The catch is baked into the concept: the chip runs exactly one model, frozen at manufacture. Want a new model? Tape out a new chip — which Taalas says takes about two months.

AMD’s stated plan is to fold the technology into its Instinct GPU and Helios rack lineup — data-centre gear, not consumer cards. But the architecture itself is the story, because of what it doesn’t need.

What Is Hardwired Inference?

What is hardwired inference? Every AI chatbot you have ever used runs the same way: the model’s parameters (its “weights”) sit in memory chips, and a processor shuttles them across to compute hardware over and over, token by token. That memory traffic is why AI needs so much HBM and DRAM, why tokens cost money, and why answers take a beat to arrive. Hardwired inference removes the trip. The weights are etched permanently into the chip as mask-ROM — the model is the circuit. Nothing is fetched, so nothing waits, and the memory bill largely disappears. The trade-off is rigidity: a chip that can only ever run the one model it was built for.

If that sounds like a curiosity, consider the price point. Taalas has talked about production inference at a fraction of a cent per million tokens — against the dollars per million that API providers charge for frontier models. For a fixed, high-demand model, “cheap and instant” is not an edge case. It is most of the traffic.

When the Impossible Happens

Every market that has ever been “essential” has a story like this one, and the pattern is always the same: the incumbents aren’t wrong about the present. They’re wrong about the future arriving from a direction nobody was watching.

In 2010, the digital camera industry shipped about 121 million units a year. Smartphones existed, but were dismissed as phones. By 2018, annual shipments had fallen to 19 million — an 84% collapse in eight years, per CIPA figures reported by Digital Camera World. By 2023, shipments were down 94% from the peak. Nikon and Canon didn’t misjudge camera quality. They misjudged that “good enough, always in your pocket” would beat “better, but a separate device.”

New York’s taxi medallions — the city-issued licences that made taxi markets closed and scarce — sold for over US$1 million each at their 2013 peak, treated by their owners as safer than a mortgage. By 2018, after Uber and Lyft had added supply that the medallion system said was impossible, a medallion was worth about $170,000. Drivers who had borrowed against them at the peak were financially ruined. The licence didn’t get worse. The assumption underneath it did.

The common thread: in each case the disruption came from technology that didn’t compete on the incumbents’ terms. Smartphones didn’t make better cameras. Ride-sharing didn’t build better taxis. Hardwired silicon doesn’t run better models — it runs one model, absurdly cheaply, with no memory in the loop at all.

The Memory Market Looks Untouchable Right Now

Which brings us to RAM. The DRAM and HBM market is currently the closest thing the technology world has to a sure thing. We have covered the boom from several angles: New Zealand retailers charging up to $850 for RAM kits that cost $150 a year earlier, Samsung’s profits multiplying as chip shortages bite into 2028, and Chinese memory maker CXMT’s Shanghai debut soaring 472% on AI memory demand. Nvidia has raised server prices 15% and pointed at memory costs. Phone, laptop and GPU buyers everywhere — including every school and household in New Zealand — are paying the AI memory premium right now.

The bull case is simple and, today, airtight: every AI query moves weights through memory, AI demand is exploding, therefore memory demand explodes with it. Industry forecasts assume the relationship between AI growth and memory growth is a law of nature.

Hardwired inference breaks exactly that link. If a meaningful share of the world’s inference runs on chips where the weights never leave the die, the per-query memory demand doesn’t just shrink — it approaches zero. The “AI eats the world’s DRAM” equation quietly stops being an equation.

Why It Probably Won’t Happen Soon (and Why That Isn’t Reassuring)

Honesty requires the caveats, and they are substantial. The HC1 runs one mid-sized model from 2024, aggressively compressed, with acknowledged quality loss. Frontier models are hundreds of times larger and don’t fit on any single die — the HC1 is already close to the practical size limit of conventional lithography, so the big-model path requires multi-chip partitioning that hasn’t been demonstrated at scale. AMD’s roadmap targets data centres, where flexibility still rules. And a fixed-function chip is a bet that a specific model stays worth running — a strange bet in an industry that ships a new frontier model every few months.

But note what every one of those caveats has in common: they are arguments about timeline, not about direction. Two months per tape-out improves. Multi-chip partitioning improves. The economics improve with every process node. And the memory market’s boom is priced on none of that ever mattering — on inference being memory-hungry forever, on the 2026 shortage arithmetic holding into 2028 and beyond.

That is precisely the gap where the camera industry lived in 2010 and the medallion owners lived in 2013. “Seems impossible” has never once stopped a technology shift. It has only ever delayed the press release announcing one.

What It Means for New Zealand

For now, the practical NZ impact runs through prices: the memory premium is already landing on phones, laptops and RAM in local stores, and it hits schools, community groups and low-income households first. The long-range twist worth watching: if fixed-function inference chips eventually absorb a chunk of AI’s memory appetite, some of the pressure bidding up RAM could ease — the same force currently squeezing NZ consumers would, ironically, be the thing that relieves it. Samsung’s record chip profits and the projected shortage into 2028 are priced on that appetite never wavering.

FAQ

Is the 17,000 tokens per second figure confirmed? It’s a vendor claim, consistent across Taalas’s launch materials and coverage by Forbes, Data Center Dynamics and others. No independent benchmark has been published. The acquisition itself is confirmed by AMD.

Can these chips run frontier models like GPT-class systems? Not on one die today. The HC1 runs an 8-billion-parameter model; frontier models are orders of magnitude larger and would need multi-chip approaches that remain unproven for this architecture.

Is a RAM market crash imminent? No. Fixed-function inference is a young, narrow technology, and AMD’s integration is aimed at data centres. The point of the history is not prediction by date — it’s that markets with record profits and universal certainty are exactly the ones technology has historically ambushed.

Why did AMD buy Taalas if the tech is so limited? For the same reason incumbents buy any capability that rewrites cost curves: the patent portfolio (Taalas has pending patents on storing weights and computing in the same structures), the team, and the option on a future where inference is specialised silicon rather than general-purpose GPUs.

🔍 THE BOTTOM LINE

The Taalas acquisition is small news with a large shadow. One chip that etches a model into silicon won’t crash the memory market. But the mechanism now exists, it is cheap, it is patented, and it is owned by one of the world’s largest chipmakers — while the memory market records profits that assume its demand curve is a permanent feature of the universe. Cameras were 84% of a market one decade and a rounding error the next. Medallions were million-dollar assets and then weren’t. The RAM boom may run for years yet. But “seems impossible until it happened” is not a prophecy — it’s just the most reliable pattern in the industry’s history.

📰 Sources


— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.

Sources: AMD IR, Data Center Dynamics, The Register, World Economic Forum, The Flaw, Taalas