ai-inference

(25)
Company
CoreWeave
Ai · Enterprise · Developer Tools

CoreWeave

The crypto-mining startup that became the picks-and-shovels supplier of the AI boom - renting Nvidia GPUs to the labs building the future, one megacluster at a time.

gpu-cloud · ai-infrastructureRead →
Company
Akamai Technologies
Enterprise · Saas · Developer Tools

Akamai Technologies

Akamai began with a pixel hidden on Disney's website. Twenty-seven years later, the network built to beat the 'World Wide Wait' is becoming a security perimeter and an AI cloud - without giving up the pipes that made it matter.

akamai · cybersecurityRead →
Company
DigitalOcean
Developer Tools · Saas · Enterprise

DigitalOcean

DigitalOcean built a public cloud by making servers feel less like a procurement exercise and more like a developer tool. Now it is trying to carry that same clarity into the expensive, unruly business of production AI.

digitalocean · cloud-computingRead →
Company
Cactus
Ai · Developer Tools · Hardware

Cactus

Cactus is a cross-platform, open-source AI inference engine built in C/C++ to run language, speech, and vision models directly on smartphones, laptops, wearables, and other low-power edge hardware. It gives mobile developers Flutter, React Native, Kotlin, and Swift SDKs to deploy quantized models on-device - cutting latency to under 100ms, keeping data private, working offline, and avoiding cloud API bills - with automatic cloud fallback for heavier tasks. A YC Summer 2025 company, Cactus already powers production apps serving 500,000+ weekly inference tasks.

cactus · cactus-computeRead →
Legend
Veerbhan Kheterpal
Founder · Executive · Engineer

Veerbhan Kheterpal

Veerbhan Kheterpal is the co-founder and CEO of Quadric, a Burlingame, California processor IP company building the Chimera general-purpose neural processor (GPNPU) for on-device AI inference. A three-time technology founder with a PhD from Carnegie Mellon University, he previously co-founded Fabbrix (acquired by PDF Solutions) and was a technical co-founder of 21, Inc. Quadric's software-first architecture aims to run any AI model without the obsolescence that plagues fixed-function AI accelerators, and the company has raised roughly $90M across its funding rounds while scaling licensing revenue sharply through 2025.

veerbhan-kheterpal · quadricRead →
Company
Neurophos
Hardware · Ai · Climate

Neurophos

Neurophos is an Austin-based photonics startup building an Optical Processing Unit (OPU) that uses light instead of electrons to run AI inference. Spun out of Duke University and the Metacept incubator, the company packs more than a million micron-scale metamaterial optical modulators - roughly 10,000x smaller than conventional photonic elements - onto a single chip to perform the matrix-vector multiplication at the heart of large AI models. It positions the OPU as an energy-efficient, drop-in alternative to GPUs for data-center inference, and raised a $110M Series A led by Gates Frontier in January 2026.

photonic-computing · optical-processing-unitRead →
Company
DEEPX
Ai · Hardware · Enterprise

DEEPX

DEEPX is a South Korean fabless semiconductor startup building ultra-low-power neural processing units (NPUs) for on-device and edge AI. Founded in 2018 by former Apple chip designer Lokwon Kim, the company designs AI accelerators - led by its 25 TOPS DX-M1 chip - that run vision and generative AI models at under 5 watts, aiming to move intelligence out of the data center and into robots, cameras, factories, mobility and everyday devices. It raised an $80.5M Series C in 2024 at a $529M valuation and positions itself as a 'Physical AI infrastructure company.'

deepx · npuRead →
Company
Parasail
Ai · Developer Tools · Enterprise

Parasail

Parasail is a San Mateo-based AI inference cloud that lets developers run open-source and custom models without buying or committing to their own GPUs. Instead of owning chips, it orchestrates rented GPU capacity across roughly 40 data centers in 15 countries plus its own clusters, automatically routing each workload to the cheapest, fastest hardware. Founded in 2023 by Mike Henry (ex-Groq CPO, Mythic founder) and Tim Harris (Swift Navigation founder), the company processes on the order of 500 billion tokens a day and pitches a token-based economy - pay for output, not hardware.

ai-inference · gpu-cloudRead →
Company
Quadric
Ai · Hardware · Developer Tools

Quadric

Quadric is a semiconductor IP company that licenses the Chimera family of general-purpose neural processing units (GPNPUs) - a single, fully C++ programmable processor core that runs machine learning inference alongside classic DSP and control code in one unified architecture. Founded in 2016 by three engineers who previously built bitcoin computing hardware, Quadric aims to fix the fragmentation of on-device AI silicon: instead of a fixed-function neural accelerator that breaks whenever a new model appears, its cores scale from 1 to 864 TOPS and can adopt new networks - including large language models up to 30 billion parameters - through a software port rather than a chip respin.

quadric · gpnpuRead →
Company
Tensordyne
Ai · Hardware · Enterprise

Tensordyne

Tensordyne is a Sunnyvale- and Munich-based AI hardware company building inference systems that use a hardware-native logarithmic number system to cut the cost and power of running large AI models. Formerly the computer-vision startup Recogni, it rebranded in September 2025 and in June 2026 unveiled Napier (TDN), a 3nm, air-cooled inference platform it claims delivers roughly 13x the throughput and 17x the efficiency of Nvidia's GB300 NVL72 rack. The company positions itself as a direct challenger to Nvidia in the market for profitable, high-throughput generative-AI inference.

ai-inference · ai-chipsRead →
Company
Inferact
Ai · Developer Tools · Enterprise

Inferact

Inferact is a Berkeley- and San Francisco-based AI infrastructure company founded by the creators and core maintainers of vLLM, the most widely used open-source engine for running large language models. The company commercializes vLLM - which at any moment powers inference on more than 400,000 GPUs worldwide - by funneling paid engineering resources back into the open project while building a next-generation commercial 'universal inference layer' meant to make AI inference cheaper and faster. Inferact launched publicly in January 2026 with a $150 million seed round at an $800 million valuation, co-led by Andreessen Horowitz and Lightspeed.

inferact · vllmRead →
Company
OpenInfer
Ai · Enterprise · Developer Tools

OpenInfer

OpenInfer is a San Mateo-based AI infrastructure startup building an 'Inference OS for the agentic era' - software that runs large AI models and agents directly on the CPUs, GPUs and NPUs enterprises already own, from edge devices to private data centers, without cloud lock-in. Founded by ex-Meta and Roblox systems engineers Behnam Bastani and Reza Nourai, its OpenInfer Engine claims 2-3x faster inference than Llama.cpp and Ollama on distilled DeepSeek models and works as a drop-in replacement for existing endpoints. The company raised an $8M seed round in February 2025 and has since shipped Jean, a private email-native agentic AI system, and orchestration layers for heterogeneous compute.

on-device-ai · edge-aiRead →
Legend
Nikola Borisov
Founder · Engineer · Executive

Nikola Borisov

Nikola Borisov is the co-founder and CEO of Deep Infra, a Palo Alto cloud company that runs open-source AI models as a low-cost, low-latency inference service. A former competitive programmer who scaled messaging backends to hundreds of millions of users at imo.im and helped build HalloApp's founding backend, he bet early that inference, not training, would become the real bottleneck of the AI era. Deep Infra now processes close to five trillion tokens a week and raised a $107M Series B in May 2026 backed by NVIDIA, 500 Global, Felicis and others.

nikola-borisov · deep-infraRead →
Legend
Kam Eshghi
Founder · Executive · Operator

Kam Eshghi

Kam Eshghi is Chief Business Officer at OpenInfer, the Inference OS for the agentic era. He co-founded Lightbits Labs, where he helped pioneer NVMe/TCP storage and drove its adoption across data centers and AI clouds before the company was acquired by NVIDIA. Earlier in his career he led strategic alliances at DSSD (acquired by EMC for $1B), built the NVMe controller business at IDT, and shaped server chipset and Ethernet switch roadmaps at Intel. MIT EECS and Berkeley Haas MBA. Forbes Technology Council member.

openinfer · lightbits-labsRead →
Company
FriendliAI
Ai · Enterprise · Developer Tools

FriendliAI

FriendliAI is a San Francisco-based AI inference cloud built by the researchers who invented continuous batching - the technique now standard across the industry. Its platform serves open-weight and custom generative AI models in production with high throughput, lower GPU costs, and 99.99% reliability, used by customers including LG and Twelve Labs.

ai-inference · llmRead →
Company
Gimlet Labs, Inc.
Ai · Enterprise · Developer Tools

Gimlet Labs, Inc.

Gimlet Labs is a San Francisco applied research company building the first multi-silicon inference cloud - software that runs AI workloads simultaneously across CPUs, GPUs, and specialized accelerators. Founded by ex-Google and ex-NVIDIA engineers behind the Pixie observability project, Gimlet ships Gimlet Cloud (serverless agent inference) and kforge (autonomous kernel generation from PyTorch). The company exited stealth in October 2025 with eight-figure revenue and raised an $80M Series A led by Menlo Ventures in March 2026.

ai-infrastructure · ai-inferenceRead →
Company
Deep Infra Inc.
Ai · Developer Tools · Enterprise

Deep Infra Inc.

Deep Infra is a Palo Alto-based AI inference cloud that lets developers run hundreds of open-source machine learning models - large language models, image and video generation, speech, and embeddings - through a simple, OpenAI-compatible, pay-per-use API. The company owns and operates its own GPU fleet across multiple US data centers, processing trillions of tokens per week for companies that want production AI without managing infrastructure.

ai-inference · gpu-cloudRead →
Company
EnCharge AI
Hardware · Ai · Enterprise

EnCharge AI

EnCharge AI is a Santa Clara semiconductor startup, spun out of Princeton University, building AI accelerator chips based on analog in-memory computing. Its charge-based architecture runs heavy AI workloads on laptops, workstations and edge devices at roughly 20x the energy efficiency of conventional GPUs. Its first product, the EN100 accelerator, delivers 200+ TOPS within an 8.25W power budget, aiming to move generative AI out of the data center and onto the devices people actually hold.

analog-in-memory-computing · ai-acceleratorRead →
Company
Panthalassa
Climate · Hardware · Ai

Panthalassa

Panthalassa is a Portland, Oregon public-benefit corporation building autonomous, mass-produced floating nodes that harvest wave energy in the open ocean and use it on the spot to run AI compute, beaming inference results back to land by satellite. Founded in 2016, the company raised a $140M Series B in May 2026 led by Peter Thiel to finish its pilot manufacturing facility and deploy its Ocean-3 node series.

ocean-energy · wave-energyRead →
Company
Etched
Ai · Hardware · Enterprise

Etched

Etched is an American semiconductor startup building Sohu, an application-specific integrated circuit (ASIC) designed to run only one thing: transformer models. By burning the transformer architecture directly into silicon, Etched claims an 8-chip Sohu server can match the throughput of 160 Nvidia H100 GPUs on Llama-70B inference. Founded in 2022 by three Harvard dropouts, the company has raised over $625M, including a recent $500M round at a reported $5B valuation.

ai-chip · transformer-asicRead →
Company
Positron AI
Ai · Hardware · Enterprise

Positron AI

Positron AI designs purpose-built inference hardware for transformer models, aiming to make Nvidia GPUs optional for running large language models at production scale. Its first product, Atlas, ships from US fabs and claims roughly 3x lower latency and 4x better performance-per-watt versus an H100 system.

ai-inference · ai-hardwareRead →
Legend
Brian Yoo
Executive · Operator · Founder

Brian Yoo

Brian Yoo is Chief Business Officer at FriendliAI, an AI inference platform that helps enterprises deploy and scale large language models faster and cheaper. Before joining FriendliAI in April 2026, he spent nearly a decade as COO at Moloco, where he scaled the AI-driven advertising company from 10 employees to 600+ and grew revenue over 500x to $250M+, helping it reach a ~$4 billion valuation. A Cornell-trained operations research engineer turned business operator, Yoo brings deep expertise in scaling AI-native companies through go-to-market strategy, financial operations, and global team building.

ai-inference · llmRead →
Legend
Nicolas Muller
Founder · Executive · Engineer

Nicolas Muller

Nicolas Muller is the co-founder and CEO of Arago, a Paris-based deep-tech startup building a photonic AI accelerator called JEF that uses light instead of transistors to perform AI inference at 10x lower energy than leading GPUs. Coming from a physics and machine learning background - including an ML degree from MIT - Muller co-founded Arago in 2024 alongside Eliott Sarrey and Ambroise Müller. The company raised an oversubscribed $26M seed round in July 2025 backed by Earlybird, Protagonist, Visionaries Tomorrow, and angels from Apple, Datadog, Hugging Face and Arm. Arago's approach - top-down from inference efficiency requirements, not bottom-up from optics research - delivered a working prototype in under 12 months.

photonic-computing · ai-hardwareRead →
Legend
Zain Asgar
Founder · Engineer · Executive

Zain Asgar

Zain Asgar is the Co-Founder and CEO of Gimlet Labs, a San Francisco-based AI infrastructure company building the world's first multi-silicon inference cloud. With a PhD from Stanford in electrical engineering focused on GPU energy modeling, Asgar previously led engineering at Google AI (where his work became Google Lens) and founded Pixie Labs, a Kubernetes-native observability platform acquired by New Relic in 2020. At Gimlet Labs, he is tackling one of AI's most pressing infrastructure challenges: making AI inference 3-10x more efficient by intelligently routing workloads across heterogeneous hardware including NVIDIA, AMD, Intel, ARM, and specialized accelerators like Cerebras.

ai-inference · heterogeneous-computingRead →
Legend
Mitesh Agrawal
Founder · Executive · Operator

Mitesh Agrawal

Mitesh Agrawal is the CEO of Positron AI, a Reno-based AI inference hardware startup that has raised $305M and achieved unicorn status in under 34 months. A co-founder and long-time COO of Lambda Labs - where he helped grow revenues from $500K to ~$500M - Agrawal brought his front-row view of GPU limitations to Positron, a company building purpose-built silicon (the Atlas system, and next-gen Asimov chip) that delivers 3x the performance of Nvidia GPUs at one-third the power. Backed by Arm, Qatar Investment Authority, Jump Trading, and others, Positron is building America's first at-scale AI inference silicon company.

ai-hardware · ceoRead →