Artificial intelligence workloads have been shifting from training large models to running them for everyday use, a process known as inference. Nvidia's pricey graphics processors dominate training, but inference has attracted specialist chipmakers hoping to serve chatbots, coding assistants and voice agents that demand very fast responses. One of those specialists, Santa Clara, California based d-Matrix, has spent years building memory-centric accelerators for exactly that kind of low-latency work.
Reuters reported on September 10, 2026 that d-Matrix said it will adopt Nvidia's NVLink Fusion chip-linking technology so its processors can be used directly inside the semiconductor giant's data-center systems. NVLink Fusion includes connectors and specialized memory so custom AI chips can plug into Nvidia's larger data-center systems, the startup said. The new chips, called Raptor, will plug into Nvidia's server racks using NVLink Fusion.
d-Matrix said in a press release dated September 10, 2026 that the collaboration includes a multi-year product roadmap giving its XPUs entry into the widely deployed Nvidia AI factory ecosystem. Raptor is expected to tape out before the end of 2026, and initial availability of Raptor XPUs integrated into Nvidia's MGX rack is expected in the fourth quarter of 2027. Nvidia's own blog, published on September 10, 2026, confirmed the news and said d-Matrix is joining a growing roster of partners building on the Nvidia AI platform.
The deal signals that Nvidia, which holds a commanding position in AI training, is willing to open its rack architecture to rival silicon aimed at inference. For d-Matrix, the arrangement offers a lower-risk path from custom silicon to large-scale deployment. For customers, it offers a way to mix specialized inference chips with Nvidia's broader AI infrastructure.
Key Facts
Raptor is d-Matrix's next-generation inference XPU, designed from the ground up for integration with NVLink Fusion and the Nvidia MGX rack-scale ecosystem. It is a follow-on to the d-Matrix Corsair XPU platform currently in production and is backed by more than 100 patents. Raptor uses a first-of-its-kind 3D DRAM stacking approach that brings a DRAM memory chip and an SRAM compute chip together into a single two-story package. Technical details were recently published by IEEE and previewed by d-Matrix co-founder and CTO Sudeep Bhoja at the 2026 Hot Chips conference.
Timeline: d-Matrix expects Raptor to tape out before the end of 2026. Initial availability of Raptor XPUs integrated into the Nvidia MGX rack is expected in the fourth quarter of 2027. The collaboration covers a multi-year product roadmap, and Raptor is being actively evaluated at AI hyperscalers and frontier labs, according to the company.
Money and investors: d-Matrix was valued at $2 billion when it raised $450 million in 2025, and it received $110 million in 2023 with Microsoft backing. Microsoft has backed the startup since that 2023 financing round. The startup shipped its first AI chip in November 2024. The financial terms of the Nvidia collaboration were not disclosed.
The d-Matrix rack, enabled by the MGX platform, will feature modular cable-free trays built with Nvidia's MGX ecosystem and supply chain. d-Matrix plans to integrate Nvidia Vera CPUs, Nvidia ConnectX-9 SuperNICs, Nvidia BlueField-4 DPUs and Nvidia Spectrum-X Ethernet networking. It is also partnering with connectivity firm Astera Labs to build custom solutions to ensure fast data flow across the system. d-Matrix plans to connect its XPUs in a single high-bandwidth, low-latency scale-up domain, and its racks can work alongside Nvidia GPU-based systems such as Nvidia Vera Rubin NVL72 for disaggregated inference.
Sid Sheth, cofounder and CEO of d-Matrix, said during a press briefing that demand for inference is soaring, but capital, time and energy remain finite. He said that with NVLink Fusion and MGX, d-Matrix can integrate its Raptor XPUs into a broadly deployed, liquid-cooled architecture, giving customers a faster, lower-risk path to deploy and scale ultralow-latency inference. In the official announcement, Sheth called the collaboration a defining moment on the company's journey to infinite inference, accessible to all. Jensen Huang, founder and CEO of Nvidia, said that NVLink Fusion enables partners to integrate custom silicon with Nvidia's deep ecosystem of NVLink, advanced packaging, rack-scale systems and networking technologies.
Analysis
The bigger picture here is that Nvidia is converting a potential competitive threat into an ecosystem dependency. Inference specialists such as d-Matrix exist because general-purpose GPUs are not always the most efficient way to run the decode phase of AI workloads, where latency and token generation speed matter enormously. By opening NVLink Fusion and MGX to partners, Nvidia keeps those specialists inside its rack, its networking and its software stack rather than letting them build stand-alone alternatives. Reuters reported on September 10, 2026 that the combined systems are aimed at fast, low-latency AI services such as coding assistants, chatbots and voice agents. Those are exactly the workloads where a memory-centric design with 3D DRAM stacking could shine.
For d-Matrix, the trade-off is strategic. The startup gets access to a broadly deployed, liquid-cooled rack architecture and a supply chain it could not easily replicate, which accelerates the path from tape-out to revenue. But the arrangement also makes d-Matrix partly dependent on Nvidia's roadmap and rack cadence, a risk mitigated somewhat by the company's own Corsair platform, its more than 100 patents and its $2 billion valuation. The startup did not disclose the financial terms of the collaboration, leaving the commercial upside unclear.
The competitive context extends beyond Nvidia. d-Matrix is partnering with Astera Labs, a connectivity specialist, to deliver custom solutions for fast data flow. Astera Labs has its own interest in the rack-scale transition, since faster, more complex racks require more connectivity silicon. The involvement of Microsoft, which has backed d-Matrix since the $110 million round in 2023, adds another layer: a major cloud customer with a stake in seeing inference costs fall. What this really means is that the AI infrastructure market is segmenting into layers, with Nvidia owning the rack, networking and platform layer while smaller firms compete on processor architectures tuned for specific phases of inference.
There are open questions. Raptor still has to tape out and then prove itself in hyperscaler evaluations. Nvidia's blog said Raptor is being actively evaluated at AI hyperscalers and frontier labs, but evaluations are not deployments. The fourth quarter of 2027 target gives competitors time to respond, and Nvidia's own next-generation GPUs, including the Vera Rubin NVL72, will keep improving. Still, the direction is clear: inference is becoming the battleground, and rack-scale integration is the weapon.
Why It Matters
The shift from training to inference is changing what data centers need. Training rewards raw throughput and massive clusters, while inference rewards low latency, energy efficiency and cost per token. d-Matrix's memory-centric Raptor design and its 3D DRAM stacking approach target those requirements directly. If Raptor works as intended, operators could use a heterogeneous disaggregation model in which Nvidia Vera Rubin GPUs handle the compute-intensive prefill phase of AI coding while d-Matrix XPUs accelerate the latency-sensitive decode phase. That kind of split could lower the cost of serving AI features at scale.
The announcement also shows how Nvidia is defending its position. NVLink Fusion extends the platform's openness to XPUs and CPUs, allowing silicon companies to focus on processor innovations while using Nvidia infrastructure to deploy them at AI factory scale. That framing positions Nvidia as the neutral platform for AI factories, even as it competes with the very partners it enables. For customers, the benefit is choice within a common rack and networking standard. For Nvidia, the benefit is that every partner rack still carries Nvidia switches, DPUs, SuperNICs and Ethernet.
Finally, the deal is a signal to other inference startups. If d-Matrix can plug into MGX and NVLink Fusion, then the barrier to entry for custom inference silicon may be lower than it appears, but the barrier to building a full rack ecosystem remains high. That dynamic is likely to push more chipmakers toward partnership rather than confrontation with Nvidia.
Next Up
The immediate milestone is Raptor's tape-out before the end of 2026. After that, d-Matrix must convert hyperscaler and frontier lab evaluations into design wins. The company has said initial availability of Raptor XPUs integrated into the Nvidia MGX rack is expected in the fourth quarter of 2027, so the next 12 to 18 months will be about silicon validation, software readiness and customer commitments. Astera Labs's custom connectivity work will also need to be proven in the rack.
Investors and competitors will watch whether the multi-year roadmap produces additional XPU generations and whether Microsoft, already a backer, deepens its involvement. Nvidia, meanwhile, will keep expanding the list of NVLink Fusion partners, a roster that now includes d-Matrix. The question is not whether inference matters, but how much of the inference market Nvidia can keep inside its own racks.
Comments (0)
Log in or sign up to leave a comment.
No comments yet. Be the first to share your thoughts.