Who decides which chip serves which part of an AI request? d-Matrix answered that question on August 3 by acquiring Wallaroo.ai, a software company that packages and orchestrates AI inference workloads. Financial terms were not disclosed.
d-Matrix’s Corsair accelerators are built to share work with GPUs: GPUs handle the compute-heavy prefill phase, while Corsair takes the token-by-token decode phase that strains memory bandwidth. Orchestrating that split across a live cluster is a software problem, and Wallaroo’s control plane is now the answer d-Matrix owns.
Wallaroo’s runtime targets x86, Arm and GPU hardware in cloud, on-premises, edge and air-gapped settings, and supports common LLM serving stacks like vLLM and SGLang. The two companies had already converged on the same architecture, describing agentic traffic as mostly short requests that occasionally hit a long prompt which stalls everything queued behind it.
The deal is d-Matrix’s second acquisition in four months, following GigaIO’s data center business in April, and comes after a $275M Series C in November 2025. Rivals are consolidating too: Qualcomm bought compiler startup Modular in July, and Nebius agreed to pay $643M for Eigen AI.
Parasail, a cloud already pairing Corsair with Nvidia Hopper and Blackwell GPUs, shows the hybrid setup in action. d-Matrix argues the next mixed-silicon rack will win or lose on how fast a model goes from evaluation to production traffic.