inference-engine

(3)
Company
Inferact
Ai · Developer Tools · Enterprise

Inferact

Inferact is a Berkeley- and San Francisco-based AI infrastructure company founded by the creators and core maintainers of vLLM, the most widely used open-source engine for running large language models. The company commercializes vLLM - which at any moment powers inference on more than 400,000 GPUs worldwide - by funneling paid engineering resources back into the open project while building a next-generation commercial 'universal inference layer' meant to make AI inference cheaper and faster. Inferact launched publicly in January 2026 with a $150 million seed round at an $800 million valuation, co-led by Andreessen Horowitz and Lightspeed.

inferact · vllmRead →
Company
OpenInfer
Ai · Enterprise · Developer Tools

OpenInfer

OpenInfer is a San Mateo-based AI infrastructure startup building an 'Inference OS for the agentic era' - software that runs large AI models and agents directly on the CPUs, GPUs and NPUs enterprises already own, from edge devices to private data centers, without cloud lock-in. Founded by ex-Meta and Roblox systems engineers Behnam Bastani and Reza Nourai, its OpenInfer Engine claims 2-3x faster inference than Llama.cpp and Ollama on distilled DeepSeek models and works as a drop-in replacement for existing endpoints. The company raised an $8M seed round in February 2025 and has since shipped Jean, a private email-native agentic AI system, and orchestration layers for heterogeneous compute.

on-device-ai · edge-aiRead →
Company
Superlinked, Inc.
Ai · Developer Tools · Enterprise

Superlinked, Inc.

Superlinked is a San Francisco-based infrastructure company building the 'vector computer' - a compute and data engineering framework that turns complex enterprise data into vector embeddings for Search, RAG, Recommendations and Analytics. It now also ships an open-source Inference Engine that lets agent teams self-host every model their agents call, claiming roughly 50x cost savings vs hosted per-token APIs.

vector-embeddings · vector-searchRead →