ai-inference

(18)
Company
CoreWeave
Ai · Enterprise · Developer Tools

CoreWeave

The crypto-mining startup that became the picks-and-shovels supplier of the AI boom - renting Nvidia GPUs to the labs building the future, one megacluster at a time.

gpu-cloud · ai-infrastructureRead →
Company
Akamai Technologies
Enterprise · Saas · Developer Tools

Akamai Technologies

Akamai began with a pixel hidden on Disney's website. Twenty-seven years later, the network built to beat the 'World Wide Wait' is becoming a security perimeter and an AI cloud - without giving up the pipes that made it matter.

akamai · cybersecurityRead →
Company
DigitalOcean
Developer Tools · Saas · Enterprise

DigitalOcean

DigitalOcean built a public cloud by making servers feel less like a procurement exercise and more like a developer tool. Now it is trying to carry that same clarity into the expensive, unruly business of production AI.

digitalocean · cloud-computingRead →
Company
Cactus
Ai · Developer Tools · Hardware

Cactus

Cactus is a cross-platform, open-source AI inference engine built in C/C++ to run language, speech, and vision models directly on smartphones, laptops, wearables, and other low-power edge hardware. It gives mobile developers Flutter, React Native, Kotlin, and Swift SDKs to deploy quantized models on-device - cutting latency to under 100ms, keeping data private, working offline, and avoiding cloud API bills - with automatic cloud fallback for heavier tasks. A YC Summer 2025 company, Cactus already powers production apps serving 500,000+ weekly inference tasks.

cactus · cactus-computeRead →
Company
Neurophos
Hardware · Ai · Climate

Neurophos

Neurophos is an Austin-based photonics startup building an Optical Processing Unit (OPU) that uses light instead of electrons to run AI inference. Spun out of Duke University and the Metacept incubator, the company packs more than a million micron-scale metamaterial optical modulators - roughly 10,000x smaller than conventional photonic elements - onto a single chip to perform the matrix-vector multiplication at the heart of large AI models. It positions the OPU as an energy-efficient, drop-in alternative to GPUs for data-center inference, and raised a $110M Series A led by Gates Frontier in January 2026.

photonic-computing · optical-processing-unitRead →
Company
DEEPX
Ai · Hardware · Enterprise

DEEPX

DEEPX is a South Korean fabless semiconductor startup building ultra-low-power neural processing units (NPUs) for on-device and edge AI. Founded in 2018 by former Apple chip designer Lokwon Kim, the company designs AI accelerators - led by its 25 TOPS DX-M1 chip - that run vision and generative AI models at under 5 watts, aiming to move intelligence out of the data center and into robots, cameras, factories, mobility and everyday devices. It raised an $80.5M Series C in 2024 at a $529M valuation and positions itself as a 'Physical AI infrastructure company.'

deepx · npuRead →
Company
Parasail
Ai · Developer Tools · Enterprise

Parasail

Parasail is a San Mateo-based AI inference cloud that lets developers run open-source and custom models without buying or committing to their own GPUs. Instead of owning chips, it orchestrates rented GPU capacity across roughly 40 data centers in 15 countries plus its own clusters, automatically routing each workload to the cheapest, fastest hardware. Founded in 2023 by Mike Henry (ex-Groq CPO, Mythic founder) and Tim Harris (Swift Navigation founder), the company processes on the order of 500 billion tokens a day and pitches a token-based economy - pay for output, not hardware.

ai-inference · gpu-cloudRead →
Company
Quadric
Ai · Hardware · Developer Tools

Quadric

Quadric is a semiconductor IP company that licenses the Chimera family of general-purpose neural processing units (GPNPUs) - a single, fully C++ programmable processor core that runs machine learning inference alongside classic DSP and control code in one unified architecture. Founded in 2016 by three engineers who previously built bitcoin computing hardware, Quadric aims to fix the fragmentation of on-device AI silicon: instead of a fixed-function neural accelerator that breaks whenever a new model appears, its cores scale from 1 to 864 TOPS and can adopt new networks - including large language models up to 30 billion parameters - through a software port rather than a chip respin.

quadric · gpnpuRead →
Company
Tensordyne
Ai · Hardware · Enterprise

Tensordyne

Tensordyne is a Sunnyvale- and Munich-based AI hardware company building inference systems that use a hardware-native logarithmic number system to cut the cost and power of running large AI models. Formerly the computer-vision startup Recogni, it rebranded in September 2025 and in June 2026 unveiled Napier (TDN), a 3nm, air-cooled inference platform it claims delivers roughly 13x the throughput and 17x the efficiency of Nvidia's GB300 NVL72 rack. The company positions itself as a direct challenger to Nvidia in the market for profitable, high-throughput generative-AI inference.

ai-inference · ai-chipsRead →
Company
Inferact
Ai · Developer Tools · Enterprise

Inferact

Inferact is a Berkeley- and San Francisco-based AI infrastructure company founded by the creators and core maintainers of vLLM, the most widely used open-source engine for running large language models. The company commercializes vLLM - which at any moment powers inference on more than 400,000 GPUs worldwide - by funneling paid engineering resources back into the open project while building a next-generation commercial 'universal inference layer' meant to make AI inference cheaper and faster. Inferact launched publicly in January 2026 with a $150 million seed round at an $800 million valuation, co-led by Andreessen Horowitz and Lightspeed.

inferact · vllmRead →
Company
OpenInfer
Ai · Enterprise · Developer Tools

OpenInfer

OpenInfer is a San Mateo-based AI infrastructure startup building an 'Inference OS for the agentic era' - software that runs large AI models and agents directly on the CPUs, GPUs and NPUs enterprises already own, from edge devices to private data centers, without cloud lock-in. Founded by ex-Meta and Roblox systems engineers Behnam Bastani and Reza Nourai, its OpenInfer Engine claims 2-3x faster inference than Llama.cpp and Ollama on distilled DeepSeek models and works as a drop-in replacement for existing endpoints. The company raised an $8M seed round in February 2025 and has since shipped Jean, a private email-native agentic AI system, and orchestration layers for heterogeneous compute.

on-device-ai · edge-aiRead →
Company
FriendliAI
Ai · Enterprise · Developer Tools

FriendliAI

FriendliAI is a San Francisco-based AI inference cloud built by the researchers who invented continuous batching - the technique now standard across the industry. Its platform serves open-weight and custom generative AI models in production with high throughput, lower GPU costs, and 99.99% reliability, used by customers including LG and Twelve Labs.

ai-inference · llmRead →
Company
Gimlet Labs, Inc.
Ai · Enterprise · Developer Tools

Gimlet Labs, Inc.

Gimlet Labs is a San Francisco applied research company building the first multi-silicon inference cloud - software that runs AI workloads simultaneously across CPUs, GPUs, and specialized accelerators. Founded by ex-Google and ex-NVIDIA engineers behind the Pixie observability project, Gimlet ships Gimlet Cloud (serverless agent inference) and kforge (autonomous kernel generation from PyTorch). The company exited stealth in October 2025 with eight-figure revenue and raised an $80M Series A led by Menlo Ventures in March 2026.

ai-infrastructure · ai-inferenceRead →
Company
Deep Infra Inc.
Ai · Developer Tools · Enterprise

Deep Infra Inc.

Deep Infra is a Palo Alto-based AI inference cloud that lets developers run hundreds of open-source machine learning models - large language models, image and video generation, speech, and embeddings - through a simple, OpenAI-compatible, pay-per-use API. The company owns and operates its own GPU fleet across multiple US data centers, processing trillions of tokens per week for companies that want production AI without managing infrastructure.

ai-inference · gpu-cloudRead →
Company
EnCharge AI
Hardware · Ai · Enterprise

EnCharge AI

EnCharge AI is a Santa Clara semiconductor startup, spun out of Princeton University, building AI accelerator chips based on analog in-memory computing. Its charge-based architecture runs heavy AI workloads on laptops, workstations and edge devices at roughly 20x the energy efficiency of conventional GPUs. Its first product, the EN100 accelerator, delivers 200+ TOPS within an 8.25W power budget, aiming to move generative AI out of the data center and onto the devices people actually hold.

analog-in-memory-computing · ai-acceleratorRead →
Company
Panthalassa
Climate · Hardware · Ai

Panthalassa

Panthalassa is a Portland, Oregon public-benefit corporation building autonomous, mass-produced floating nodes that harvest wave energy in the open ocean and use it on the spot to run AI compute, beaming inference results back to land by satellite. Founded in 2016, the company raised a $140M Series B in May 2026 led by Peter Thiel to finish its pilot manufacturing facility and deploy its Ocean-3 node series.

ocean-energy · wave-energyRead →
Company
Etched
Ai · Hardware · Enterprise

Etched

Etched is an American semiconductor startup building Sohu, an application-specific integrated circuit (ASIC) designed to run only one thing: transformer models. By burning the transformer architecture directly into silicon, Etched claims an 8-chip Sohu server can match the throughput of 160 Nvidia H100 GPUs on Llama-70B inference. Founded in 2022 by three Harvard dropouts, the company has raised over $625M, including a recent $500M round at a reported $5B valuation.

ai-chip · transformer-asicRead →
Company
Positron AI
Ai · Hardware · Enterprise

Positron AI

Positron AI designs purpose-built inference hardware for transformer models, aiming to make Nvidia GPUs optional for running large language models at production scale. Its first product, Atlas, ships from US fabs and claims roughly 3x lower latency and 4x better performance-per-watt versus an H100 system.

ai-inference · ai-hardwareRead →