d-Matrix links Raptor XPU to NVIDIA rack architecture

d-Matrix links Raptor XPU to NVIDIA rack architecture

d-Matrix is integrating Raptor XPUs with NVIDIA NVLink Fusion infrastructure. The inference processor will enter MGX rack architecture alongside NVIDIA compute, networking, and data-processing components.


IN Brief:

  • Raptor XPUs will integrate with NVIDIA MGX rack architecture through NVLink Fusion and NVIDIA networking.
  • The processor combines a DRAM memory chip and SRAM compute chip in a stacked memory-centric architecture.
  • Tape-out is expected before year-end, with initial integrated MGX systems targeted for the fourth quarter of 2027.

d-Matrix is integrating its forthcoming Raptor inference XPU with NVIDIA MGX rack-scale infrastructure using NVLink Fusion, giving the custom accelerator access to the same wider architecture used around NVIDIA CPUs, networking, data-processing units, and scale-up interconnect. The collaboration covers a multi-year roadmap and is aimed specifically at latency-sensitive AI inference.

The initial configuration will combine Raptor XPUs with NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet networking. Astera Labs is developing custom connectivity intended to maintain high-throughput data paths across the resulting mixed-processor system.

NVLink Fusion gives custom silicon suppliers a route into NVIDIA’s scale-up fabric rather than requiring each accelerator company to create an independent rack interconnect. For d-Matrix, that reduces some of the engineering work around deploying its processor as part of a complete rack rather than as an isolated accelerator card.

That distinction has become increasingly important as AI hardware scales. Processor architecture remains central, but a deployed accelerator also depends on high-speed interconnect, host processing, network adapters, management hardware, power delivery, cooling, cable routing, mechanical packaging, and software that can distribute workloads across the available compute resources.

Raptor extends d-Matrix’s existing memory-centric design approach. The company says the processor uses a three-dimensional stacking arrangement that combines a DRAM memory chip with an SRAM compute chip in a single two-level package. The architecture is intended to reduce the distance that inference data travels between storage and compute resources.

Data movement is particularly significant during generative AI inference. Prompt processing, or prefill, benefits from large amounts of parallel compute, while token-by-token decode work is sensitive to memory access and latency. d-Matrix is targeting mixed systems in which GPUs can perform prefill while Raptor XPUs handle latency-sensitive decode workloads.

Splitting the workload only works when the interfaces between processors are sufficiently fast and predictable. A gain at the accelerator can be lost if data has to cross a congested or high-latency link before the next stage can begin, so memory bandwidth, scale-up interconnect, and rack networking have to be considered alongside the XPU itself.

The MGX integration gives d-Matrix access to an existing mechanical and supply-chain platform as well as an interconnect. The company says the rack will use modular cable-free trays drawn from the MGX ecosystem, reducing the amount of server-level infrastructure that has to be engineered specifically for Raptor.

MediaTek is pursuing a separate NVLink Fusion programme around custom accelerator design, advanced packaging, and HBM integration. The d-Matrix development differs in that Raptor is its own inference-focused processor and is being designed directly into an NVIDIA MGX rack configuration.

The hardware has not yet reached silicon. d-Matrix expects Raptor to tape out before the end of 2026, after which fabrication, package assembly, bring-up, validation, software integration, and system qualification still have to be completed. Those stages will determine whether the architectural claims survive the transition into deployable hardware.

The company says Raptor is already being evaluated by hyperscalers and frontier AI laboratories and is backed by more than 100 patents. Initial availability of Raptor XPUs integrated into NVIDIA MGX racks is targeted for the fourth quarter of 2027.

The schedule leaves more than a year between the current integration announcement and planned system availability, which is realistic for hardware that combines a new accelerator, three-dimensional memory integration, high-speed scale-up connectivity, and a complete rack implementation. Each part has to reach production maturity at roughly the same time.

Raptor’s case will ultimately be decided by measurements rather than topology diagrams: decode latency, sustained throughput, memory utilisation, power, and rack-level efficiency will determine whether adding a dedicated inference XPU provides enough benefit to justify another processor architecture within the AI system.


Stories for you


  • Singular launches multimodal Litavis SPAD image sensor

    Singular launches multimodal Litavis SPAD image sensor

    Singular Photonics has launched Litavis, a programmable multimodal SPAD sensor. The Edinburgh-developed device combines photon counting, timing, histogramming, and in-pixel processing within one configurable imaging platform.


  • NVIDIA applies sovereign AI to hardware supply chains

    NVIDIA applies sovereign AI to hardware supply chains

    NVIDIA and Palantir are applying sovereign AI to supply chains. The initial deployment targets materials allocation and constraint management across NVIDIA’s complex AI hardware manufacturing network.