# Inferact

> Inferact is a Berkeley- and San Francisco-based AI infrastructure company founded by the creators and core maintainers of vLLM, the most widely used open-source engine for running large language models. The company commercializes vLLM - which at any moment powers inference on more than 400,000 GPUs worldwide - by funneling paid engineering resources back into the open project while building a next-generation commercial 'universal inference layer' meant to make AI inference cheaper and faster. Inferact launched publicly in January 2026 with a $150 million seed round at an $800 million valuation, co-led by Andreessen Horowitz and Lightspeed.

- **Founded:** 2025
- **Headquarters:** San Francisco, California, United States
- **Founders:** Simon Mo (Co-Founder & CEO), Woosuk Kwon (Co-Founder), Kaichao You (Co-Founder), Roger Wang (Co-Founder)
- **Team size:** ~22 employees
- **Products:** vLLM, Universal Inference Layer (in development)
- **Notable:** Raised a $150M seed round - unusually large for a seed - at an $800M valuation in January 2026., Built vLLM into the most widely adopted open-source LLM inference engine, running on 400,000+ GPUs concurrently worldwide., Attracted both a16z and Lightspeed as co-leads, with Sequoia, Altimeter, Redpoint, and ZhenFund participating.

## Products & services

- **vLLM** — The open-source large language model inference and serving engine created by the founders. It manages continuous batching, KV caching, and low-level chip optimization so that diverse models run efficiently across varied hardware. At any moment it runs on 400,000+ GPUs worldwide, with 2,000+ contributors and 50+ core developers.
- **Universal Inference Layer (in development)** — A planned next-generation commercial inference engine, built on the vLLM foundation, intended to serve any model on any hardware at lower cost while collaborating with existing inference providers rather than competing with them.

## Achievements

- Raised a $150M seed round - unusually large for a seed - at an $800M valuation in January 2026.
- Built vLLM into the most widely adopted open-source LLM inference engine, running on 400,000+ GPUs concurrently worldwide.
- Attracted both a16z and Lightspeed as co-leads, with Sequoia, Altimeter, Redpoint, and ZhenFund participating.
- vLLM adopted in production by Meta, Google, Character.ai, and Amazon.

## Latest updates

- **2026-01** — Inferact launched publicly and announced a $150M seed round at an $800M valuation, co-led by a16z and Lightspeed.
- **2026-01** — TechCrunch reported that vLLM's existing users include Amazon's cloud service and the Amazon Shopping app; competitor SGLang commercialized as RadixArk.

## Links

- Website: https://inferact.ai
- LinkedIn: https://www.linkedin.com/company/inferact
- Twitter/X: https://x.com/inferact
- GitHub: https://github.com/vllm-project/vllm

---

Profile page: https://yespress.io/inferact
Published by YesPress — https://yespress.io
Last updated: 2026-08-04
