September 2026 · NVIDIA NVLink Fusion collaborationAugust 2026 · Wallaroo.ai acquisitionJuly 2026 · Parasail deploymentJune 2026 · Corsair full productionSeptember 2026 · NVIDIA NVLink Fusion collaborationAugust 2026 · Wallaroo.ai acquisitionJuly 2026 · Parasail deploymentJune 2026 · Corsair full production

Company / AI infrastructure

The Chip Was Only the Beginning

d-Matrix bet early that serving AI answers would become its own computing problem. Seven years later, the Santa Clara company is building the chips, network links, rack designs and software to solve it.

In 2019, the fashionable problem in artificial intelligence was training a model. Sid Sheth and Sudeep Bhoja founded d-Matrix around a less glamorous question: what happens after the training ends and millions of people start asking that model for answers? A training run can take its time. A person waiting for a voice agent, coding assistant or chatbot cannot. The distinction sounds small until a data center has to repeat that answer, token by token, all day.

The short version
  • d-Matrix makes AI inference hardware for data centers, led by the Corsair accelerator.
  • Its approach places computation close to memory to reduce the time and energy spent moving model data.
  • The company has added networking, rack designs and deployment software because a fast chip alone cannot make a fast service.
  • Its public commercial test includes Parasail, which is pairing Corsair with NVIDIA GPUs.

The bet turned out to have a second act. Generative AI brought inference into the foreground, but it also made clear that speed depends on more than arithmetic. Model weights have to travel through memory. Results have to cross cards, servers and racks. Software has to place work where it belongs. d-Matrix began as a semiconductor company and gradually collected the rest of those chores.

d-Matrix co-founders Sid Sheth and Sudeep Bhoja
Sid Sheth and Sudeep Bhoja founded the company before the chatbot became everyone’s favorite stopwatch.

The expensive journey of a number

A large language model generates an answer by repeatedly reading parameters and producing the next token. Conventional processors keep much of their working memory apart from their compute units. Moving numbers back and forth costs time and power. d-Matrix’s digital in-memory compute design brings computation into the memory architecture itself. Corsair, its flagship PCIe accelerator, packages that idea into chiplets intended for high-throughput, low-latency inference.

That does not mean every AI task should leave a GPU. There are two broad phases in many language-model requests: prefill processes the incoming prompt, while decode generates the answer. Prefill can be compute-heavy. Decode can be more sensitive to memory bandwidth and latency. The increasingly specific d-Matrix proposition is to let different processors do the work that suits them. A GPU may handle prefill; Corsair may take decode. For a cloud operator with GPUs already installed, that is a more practical pitch than replacing the fleet outright.

A request, divided by its bottlenecks
01 · PrefillProcess the prompt on compute-rich hardware, often a GPU.
02 · DecodeGenerate tokens on Corsair where memory movement matters.
03 · ServeConnect cards and racks, then deliver the answer to the user.
The split is workload dependent. The useful metric is the cost and response time of the whole request.

The company’s performance figures should be read in that light. d-Matrix says its systems can generate up to 30,000 tokens a second at 2 milliseconds per token on a Llama 70B model. Its Parasail announcement says a mixed GPU-Corsair service could deliver up to ten times faster token generation on selected workloads. Both are company claims tied to configurations and workload choices. They are interesting precisely because inference buyers have to check the same numbers against their own model, batch size, prompt length and quality requirements.

$275mSeries C, November 2025Company announcement
2026Corsair enters full productionJune announcement
2Acquisitions in four monthsGigaIO assets and Wallaroo.ai

A card cannot outrun its rack

Corsair was unveiled in 2024. Then came JetStream, a network interface card designed to move data between accelerators, and SquadRack, a rack-scale reference design built with infrastructure from Arista, Broadcom and Supermicro. It is a revealing sequence. If an accelerator saves microseconds inside one server but waits on the link to another, its headline speed vanishes before the user sees it. The company’s own engineering account says the next bottleneck after compute and memory was I/O.

JetStream uses standard Ethernet and PCIe form factors, which matter to buyers because data centers have already paid for those conventions. SquadRack offers configurations from a starter server to a full rack. Aviator, the company’s software stack, converts models from familiar frameworks and runs them across multiple cards. Its published technical material describes PyTorch, MLIR and Triton DSL support. That breadth is part of the sales pitch: an AI team should be able to bring a trained model over without inventing an entirely new workflow.

d-Matrix Corsair inference accelerator card
Corsair looks like a card. Its job description now reaches across a room full of them.
d-Matrix SquadRack rack-scale inference system rendering
The rack is where the chip’s promises meet cables, switches and everyone else’s deadlines.
“Inference is bigger than any one chip. It’s now a systems problem.”
Sid Sheth, announcing the GigaIO deal in April 2026

That line explains the purchases. In April 2026, d-Matrix acquired GigaIO’s data-center business, including rack engineering talent and PCIe fabric technology. In August, it bought Wallaroo.ai’s deployment and orchestration platform. One addition addresses how the hardware fits together; the other addresses how a customer gets models running on a mixed fleet. Neither acquisition price was disclosed. They followed the company’s $275 million Series C in November 2025, which d-Matrix said valued it at $2 billion and brought its total funding to $450 million.

The customer has a GPU already

Parasail provides the clearest public example of what d-Matrix wants a buyer to do. The inference cloud announced in July 2026 that it was deploying Corsair beside NVIDIA Hopper and Blackwell GPUs. The planned division is straightforward: GPUs for compute-heavy prefill, Corsair for latency-sensitive decode, with software routing requests across the fleet. Parasail said it expected to improve the economics of selected workloads and extend the usefulness of hardware it already owned. The announcement described a deployment, with detailed case studies to follow; it was not a published independent benchmark of every possible model.

Gimlet Labs also announced plans to put Corsair alongside GPUs in its cloud for agentic inference. These are the buyers d-Matrix is courting: inference clouds, hyperscalers, frontier labs and enterprises whose models are busy enough for latency, electricity and cost per token to become serious budget lines. The company sells through direct evaluations and OEM partners. Public list prices are absent, so the actual purchase cost has to be negotiated. The easier number to see is its capital cost as a business: $110 million raised in a 2023 Series B, then $275 million in 2025, before Corsair reached full production in June 2026.

What a buyer can borrow

Profile the whole request before choosing hardware. Separate prefill from decode, measure tokens per second alongside time per token, and include networking and deployment effort in the calculation. A specialist accelerator helps when the workload has enough memory-sensitive decode and enough volume to repay a more complex fleet. For small, irregular workloads, a general-purpose GPU service may be simpler.

The company’s latest move is almost comic in its neatness: the specialist that once seemed to be challenging the GPU is now collaborating with NVIDIA. In September 2026, d-Matrix announced plans to connect its next-generation Raptor accelerator to NVIDIA MGX racks through NVLink Fusion. Raptor is planned around a two-story package stacking DRAM with SRAM compute; the company expected it to tape out before the end of 2026. That makes Raptor a roadmap, not a shipping product. Yet the partnership says something concrete about the market. The winning data center may be a mix of processors, with the right one assigned to each stage.

d-Matrix’s original wager was that AI would need an answer machine, not merely a learning machine. Its harder discovery was that an answer machine has no single part. The chip matters. So do the links between chips, the rack holding them and the software that convinces a model to run there. Inference is a race measured in milliseconds, but it is won or lost by the entire route.