Modal is a serverless cloud platform that lets developers run compute-intensive AI and data workloads - inference, training, fine-tuning, batch jobs, and sandboxed code - without managing infrastructure. Built from scratch in Rust with a custom container runtime, file system, and GPU memory snapshotting, it delivers sub-second cold starts and elastic autoscaling from zero to thousands of GPUs, billed by the second. Founded in 2021 by Erik Bernhardsson and Akshat Bubna, the New York-based company reached roughly $300M in annualized revenue and raised a $355M Series C in 2026 at a $4.65B valuation.
Enfabrica is a Silicon Valley chip startup building the networking and memory plumbing that lets tens of thousands of GPUs behave like one giant computer. Its Accelerated Compute Fabric SuperNIC (ACF-S) moves data at 3.2 terabits per second, and its EMFASYS memory fabric lets AI clusters borrow cheap DDR5 DRAM over Ethernet instead of buying more GPUs for capacity. Founded in 2019 by ex-Broadcom and Google networking engineers, the company raised roughly $290M before Nvidia paid over $900M in September 2025 to license its technology and hire CEO Rochan Sankar and much of his team.
Parasail is a San Mateo-based AI inference cloud that lets developers run open-source and custom models without buying or committing to their own GPUs. Instead of owning chips, it orchestrates rented GPU capacity across roughly 40 data centers in 15 countries plus its own clusters, automatically routing each workload to the cheapest, fastest hardware. Founded in 2023 by Mike Henry (ex-Groq CPO, Mythic founder) and Tim Harris (Swift Navigation founder), the company processes on the order of 500 billion tokens a day and pitches a token-based economy - pay for output, not hardware.
Quadric is a semiconductor IP company that licenses the Chimera family of general-purpose neural processing units (GPNPUs) - a single, fully C++ programmable processor core that runs machine learning inference alongside classic DSP and control code in one unified architecture. Founded in 2016 by three engineers who previously built bitcoin computing hardware, Quadric aims to fix the fragmentation of on-device AI silicon: instead of a fixed-function neural accelerator that breaks whenever a new model appears, its cores scale from 1 to 864 TOPS and can adopt new networks - including large language models up to 30 billion parameters - through a software port rather than a chip respin.
Nikola Borisov is the co-founder and CEO of Deep Infra, a Palo Alto cloud company that runs open-source AI models as a low-cost, low-latency inference service. A former competitive programmer who scaled messaging backends to hundreds of millions of users at imo.im and helped build HalloApp's founding backend, he bet early that inference, not training, would become the real bottleneck of the AI era. Deep Infra now processes close to five trillion tokens a week and raised a $107M Series B in May 2026 backed by NVIDIA, 500 Global, Felicis and others.
Simon Mo is the CEO and co-founder of Inferact, the startup commercializing vLLM, the open-source inference engine he helped create and now leads as lead maintainer. A PhD student at UC Berkeley's Sky Computing Lab advised by Ion Stoica and Joseph Gonzalez, Mo has spent roughly eight years building high-throughput, memory-efficient model-serving systems, from Ray Serve at Anyscale to vLLM, which now powers inference for companies including Amazon. In January 2026 he and his vLLM co-maintainers raised a $150M seed round at an $800M valuation, co-led by Andreessen Horowitz and Lightspeed.
Ying Sheng is the co-founder and CEO of RadixArk, an AI infrastructure company that spun out of SGLang, the open-source inference engine she helped create in 2023. SGLang now runs across hundreds of thousands of GPUs and serves trillions of tokens a day for Google, Microsoft, NVIDIA, xAI and others. A Stanford computer science PhD and former xAI technical staff who co-led the inference team behind Grok, she launched RadixArk in May 2026 with $100 million in seed funding led by Accel at a $400 million valuation, on a mission to make frontier-level AI infrastructure open and accessible to everyone.
Alex Yeh is the Founder and CEO of GMI Cloud, a GPU-native AI cloud infrastructure company he built from Bitcoin mining data centers into a global AI infrastructure leader in just 30 days. GMI Cloud — one of only 6 NVIDIA Reference Platform Partners worldwide — raised $82M in Series A funding in 2024 and is behind a $12 billion sovereign AI infrastructure initiative in Japan. Yeh's mission: make building AI applications as simple as building a website on Shopify.
Positron AI designs purpose-built inference hardware for transformer models, aiming to make Nvidia GPUs optional for running large language models at production scale. Its first product, Atlas, ships from US fabs and claims roughly 3x lower latency and 4x better performance-per-watt versus an H100 system.

Mitesh Agrawal is the CEO of Positron AI, a Reno-based AI inference hardware startup that has raised $305M and achieved unicorn status in under 34 months. A co-founder and long-time COO of Lambda Labs - where he helped grow revenues from $500K to ~$500M - Agrawal brought his front-row view of GPU limitations to Positron, a company building purpose-built silicon (the Atlas system, and next-gen Asimov chip) that delivers 3x the performance of Nvidia GPUs at one-third the power. Backed by Arm, Qatar Investment Authority, Jump Trading, and others, Positron is building America's first at-scale AI inference silicon company.