Close-up of a liquid-cooled AI server rack with glowing connection pathways between circuit boards, deep blue and amber lighting, photographed in a modern data centre aisle
News

d-Matrix Is Building Its Inference Chips Inside Nvidia's Racks — and That's the Point

d-Matrix's Raptor XPU — 100 TB/s of memory bandwidth from stacked 3D DRAM — will ship in Nvidia's standard liquid-cooled racks. Nvidia wins either way.

d-MatrixNvidiaNVLink FusionAI inferencesemiconductors

d-Matrix, the AI inference chip startup backed by more than $500 million in funding including Microsoft’s M12 venture arm, announced on Thursday that its next-generation Raptor processors will connect to Nvidia’s AI infrastructure through NVLink Fusion — Nvidia’s programme for letting third-party chips plug into its rack-scale systems. The announcement on Nvidia’s blog puts d-Matrix alongside Qualcomm, Arm, Marvell, MediaTek, Fujitsu, AWS and Intel as licensees of the interconnect Nvidia began licensing last year.

🔍 THE BOTTOM LINE

The company whose chips are designed to undercut Nvidia on inference costs is building its hardware to live inside Nvidia’s own racks, run on Nvidia’s networking, and deploy through Nvidia’s datacentre footprint. It looks like a truce; functionally it is a toll booth. Every Raptor deployment pays Nvidia for the interconnect, the switching, the networking and the rack — whether or not an Nvidia GPU is in it.

Why inference is the battleground

Training gets the headlines; inference pays the bills. Every token a chatbot or agent generates has to be computed, and memory bandwidth — how fast the chip can read the model’s weights — is the bottleneck that sets both the speed and the cost of that generation. Nvidia’s response to that bottleneck was to spend $20 billion last December licensing Groq’s SRAM-heavy inference architecture and hiring most of its engineers. d-Matrix has been betting on a different route to the same destination: compute logic bonded directly on top of stacked DRAM.

At the Hot Chips conference in late August, d-Matrix revealed that each Raptor card would carry 32 GB of 3D-stacked DRAM delivering 100 TB/s of memory bandwidth — by The Register’s calculation, roughly 4.5 times the memory bandwidth of Nvidia’s Rubin GPU. Where Nvidia’s Groq 3 LPX racks might need more than 2,000 LPUs to serve a trillion-parameter model at 8-bit precision, The Register calculates a single d-Matrix system would need about 64 — or 32 at 4-bit. Fewer chips per model is the pitch; it is also, per the company’s own benchmarks shown at Hot Chips, enough to serve Z.ai’s GLM 5.2 at about 3,000 tokens per second per user, according to The Next Platform.

We have tracked the inference cost war closely this year: GLM 5.2 on AMD’s MI355X already matched 80% of Blackwell throughput at half the cost, and Cerebras keeps arguing wafer-scale silicon beats GPUs on latency. The silicon alternatives are real. What none of them had was a distribution channel — the racks, cooling, networking and supply chain that make a chip deployable at AI-factory scale.

The deal has three concrete parts, per the announcement and press coverage:

  • Scale-up: Raptor XPUs will sit in Nvidia’s NVLink compute tray inside the standard NVL144 MGX rack — the same liquid-cooled architecture used for Vera Rubin systems. d-Matrix expects to offer systems with up to 144 Raptor accelerators on a single all-to-all NVLink fabric by the end of 2027, with about 2.3 TB of aggregate 3D-DRAM and roughly 7.2 PB/s of memory bandwidth per rack.
  • Scale-out: d-Matrix plans to pair its chips with Nvidia’s Vera CPUs, ConnectX-9 SuperNICs, BlueField-4 DPUs and Spectrum-X Ethernet — meaning Nvidia sells networking and DPUs even when its GPUs sit the rack out.
  • Hybrid deployment: d-Matrix racks are designed to work alongside Nvidia GPU racks for disaggregated inference — GPUs handle the compute-heavy prompt-processing phase, Raptor handles the memory-bandwidth-bound token generation.

“Raptor needs a house to put it in,” d-Matrix CEO Sid Sheth said, per The Next Platform — and while the company could have built its own liquid-cooled rack, Nvidia’s was already deployed across the datacentres its customers actually use. “We take the Raptor trays, plug them into the same NVL144 MGX rack architecture, which is widely deployed across many datacenters, and we get instant access to those datacenters.”

Raptor itself is scheduled to tape out by the end of this year and reach the market in the fourth quarter of 2027, built on a 4nm TSMC compute die fused atop a custom DRAM die.

The strategic read: Nvidia’s walled garden has an open gate — with a toll booth

Nvidia’s Q2 FY2027 earnings call put the revenue opportunity for its AI-factory platform at about $40 billion per gigawatt for Vera Rubin systems, up from roughly $18 billion per gigawatt in the Grace-Hopper era. As The Register notes, NVLink Fusion adopters like d-Matrix still buy Nvidia’s interconnect, NVSwitch, NICs, DPUs and Ethernet gear — “Nvidia stands to make a lot of money even if it’s not selling GPUs.”

Recent NVLink Fusion deals came with Nvidia writing cheques: a $3.5 billion investment in MediaTek last week and $2 billion into Marvell earlier. Neither Nvidia nor d-Matrix has disclosed whether the d-Matrix deal included investment terms.

The counterargument Nvidia would make: openness is the product. The NVIDIA blog frames NVLink Fusion as letting silicon companies focus on their processors while Nvidia supplies the proven platform — a “vertically integrated and horizontally open” stack that supports Arm, x86 and RISC-V CPUs. Both readings can be true at once. The same architecture that lets a startup avoid building rack infrastructure from scratch also makes Nvidia’s ecosystem the default substrate for every would-be challenger — which is precisely the concern we examined when Google started using Nvidia’s own playbook against it.

For inference buyers, the practical consequence is optionality without exit: you can run a specialist decode chip inside an Nvidia-shaped datacentre, and swap workloads between them, but the racks, the fabric and the software stack still orbit one vendor. Meanwhile, the U.S. Justice Department is investigating whether Nvidia structured its Groq licensing deal to avoid antitrust scrutiny — per a New York Times report carried by Reuters — which makes the pattern of paying competitors to join the platform worth watching closely.

❓ FAQ

What is d-Matrix? A Silicon Valley inference-chip startup founded in 2019, backed with more than $500 million including Microsoft’s M12. Its current Corsair product pairs with GPUs for the token-generation phase of AI inference; Raptor is the next-generation XPU detailed at Hot Chips 2026, due to tape out this year and ship in Q4 2027.

What is NVLink Fusion? Nvidia’s licensing programme that lets third-party XPUs and CPUs connect to Nvidia’s NVLink scale-up fabric, MGX rack designs and broader AI-factory platform — compute, networking, cooling and software. Partners include AWS, Arm, Intel, Fujitsu, MediaTek, Marvell, Qualcomm, Samsung and now d-Matrix.

Why does memory bandwidth matter for inference? Generating each token requires reading the model’s parameters. The faster the chip can move data from memory to compute, the faster and cheaper each token is. Raptor’s 3D-stacked DRAM design targets roughly 100 TB/s per card — SRAM-class bandwidth with DRAM-class capacity.

Does this mean d-Matrix is no longer competing with Nvidia? It still competes for the inference workload; it no longer competes on rack infrastructure. The company’s chips will deploy in Nvidia’s standard racks and pair with Nvidia CPUs, NICs and networking. Whether that is pragmatic distribution or strategic capture depends on which analyst you read — both interpretations are in play.

🔍 THE BOTTOM LINE

d-Matrix gets the thing money can’t quickly buy — a slot in the world’s dominant AI rack ecosystem. Nvidia gets a toll booth on the road into it, plus one more competitor folded into the platform. The chips that challenge Nvidia will increasingly be deployed inside Nvidia’s architecture, on Nvidia’s terms. That is not the death of chip competition, but it is a reminder that in this market, the racks are the real estate — and Nvidia owns the ground.

📰 Sources

  • NVIDIA blog (September 10, 2026)
  • The Register (September 10, 2026)
  • The Next Platform (September 10, 2026)
  • Reuters / The New York Times (September 10, 2026, DOJ/Groq reporting)
Sources: NVIDIA blog, 'd-Matrix Adopts NVIDIA NVLink Fusion for Rack-Scale XPU Deployment' (10 September 2026), The Register, 'D-Matrix drinks the Nvidia Kool-Aid with NVLink Fusion and MGX rack designs' (10 September 2026), The Next Platform, 'Startup d-Matrix Will Pair Its Raptor Memory-Based XPU To Nvidia Rackscale Iron' (10 September 2026)