AMD used its first-ever IFA opening keynote — delivered by Jack Huynh, SVP of the Computing and Graphics Group, in Berlin on 4 September 2026 — to unveil the Threadripper Halo Station, a liquid-cooled AI workstation built around the company’s fastest desktop processor and its datacentre accelerators. The pitch was blunt: the most powerful workstation on the market, and the machine that runs trillion-parameter models without renting anyone else’s cloud.
What’s in the box
The Halo Station centres on the Ryzen Threadripper PRO 9995WX — 96 Zen 5 cores, 192 threads, boost to 5.4GHz, 384MB of L3 cache, and a 350W TDP. The chip alone runs roughly US$11,000 to $12,000. Around it: 2TB of DDR5 across eight memory channels, 128 PCIe 5.0 lanes, and two Instinct MI350P accelerators in the configuration shown at IFA.
Each MI350P is a CDNA 4 card built on TSMC N3 with 144GB of HBM3E and up to 4TB/s of memory bandwidth. That’s 288GB of accelerator memory today, and AMD says there’s a “path to four” — a future configuration that would push the total to 576GB of HBM3E, with each card drawing up to 600W. At that point the liquid cooling stops being a luxury. It’s the only way the machine stays upright.
Why it matters — and why I’m only half convinced
The obvious read is that AMD is going after the desk-side AI market that Nvidia has owned since the DGX Station era. Nvidia’s answer, the GB300-based XpertStation from MSI, ships with 748GB of memory and a roughly US$100,000 price. AMD hasn’t set a price or release date for the Halo Station, but the component maths makes six figures a reasonable guess: the CPU is ~$11,000, the two MI350Ps are estimated around $20,000 each, and 2TB of DDR5 costs about $50,000 on its own right now.
What stands out here is who this is for. Not hobbyists, not most startups — but there is a real audience between “gaming PC” and “rent a cloud pod”: research labs, film studios doing on-prem inference, New Zealand government agencies and universities that can’t ship sensitive data offshore. We’ve tracked this local-inference trend in AMD’s Ryzen AI Max 395 and the GLM-5.2 on MI355X inference cost war and can’t justify a $334,000 Lenovo ThinkStation P8. A desk-side machine that runs large open-weight models on hardware you physically own is a meaningful option for exactly those buyers. AMD has been making this argument all year — it also fronts the smaller Ryzen AI Max platform for local model inference — and the Halo Station is the top of that wedge. It also just committed US$5 billion to Anthropic in a gigawatt-scale MI450 deal, so the workstation and the datacentre story are part of one push.
The half-I’m-not-convinced part: memory bandwidth. HBM3E at 4TB/s per GPU is fast, but a system that leans on 2TB of DDR5 for model weights beyond what fits in 288GB will feel the CPU-GPU bottleneck hard. Trillion-parameter models are the marketing line; the realistic sweet spot is mid-size open-weight models that fit entirely in HBM. That’s still a legitimate use case — just not the one on the slide.
The comparison that will decide it
AMD didn’t name Nvidia on stage, but the shadow contest is obvious. Nvidia’s Blackwell workstation line has more memory today (748GB in the MSI machine) and a mature software stack in CUDA. AMD’s counter is memory bandwidth per dollar and an open ecosystem — ROCm has improved enough that recent open-weight model releases run on Instinct cards on day one, which wasn’t true two years ago.
For buyers, the honest summary is: if you’re already in the CUDA ecosystem, nothing changes this week. If you’re building fresh and your models fit in 288–576GB, AMD just put a credible alternative on the desk.
FAQ
When was the AMD Threadripper Halo Station announced? 4 September 2026, in AMD’s IFA opening keynote in Berlin.
How much GPU memory does it have? Two MI350P accelerators today = 288GB of HBM3E. AMD plans support for four, which would reach 576GB.
How much will it cost? No price announced. Component estimates — an ~$11,000 CPU, ~$20,000 per accelerator, ~$50,000 for 2TB of DDR5 — point to a six-figure system.
Can it run trillion-parameter models? AMD says so in the four-GPU configuration. In practice, models that fit entirely in the HBM3E pool will perform best; spilling into system memory costs bandwidth.
New Zealand angle
Local inference without sending data offshore is a live issue for NZ organisations handling health or government data. A workstation class like this — if it lands at a sane price — is exactly the tier that would let a NZ firm run open-weight models on-prem instead of renting US cloud capacity.
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.