Racks built around d-Matrix’s Raptor inference chips should be available integrated into NVIDIA systems by the fourth quarter of 2027, the two companies said on September 10, under a multi-year roadmap that routes the silicon through NVLink Fusion and the MGX rack architecture.
The division of work is the interesting part. Rather than replacing GPUs, d-Matrix expects operators to split a single inference job by stage, running the compute-heavy prefill phase on NVIDIA accelerators and the latency-sensitive decode phase on its own XPUs. Raptor racks are designed to sit beside GPU systems including Vera Rubin NVL72.
Applications in the company’s sights are the ones where speed is worth money: coding assistants, real-time chatbots, voice agents. The trays are modular and cable-free, built through the MGX supply chain, with custom connectivity supplied by Astera Labs.
NVIDIA’s numbers for the platform put XPU-to-XPU latency three times lower than commodity Ethernet, packet rates 10 times higher, and all-to-all bandwidth at 3 TB per second per XPU over sixth-generation NVLink.
Chief executive Sid Sheth framed the decision around limited capital, time and energy, arguing the integration lowers the risk of reaching large-scale deployment. Jensen Huang’s statement cast the same deal from NVIDIA’s side, describing a path for partners to attach custom silicon while giving customers more accelerator choice.