# OpenInfer

> OpenInfer is a San Mateo-based AI infrastructure startup building an 'Inference OS for the agentic era' - software that runs large AI models and agents directly on the CPUs, GPUs and NPUs enterprises already own, from edge devices to private data centers, without cloud lock-in. Founded by ex-Meta and Roblox systems engineers Behnam Bastani and Reza Nourai, its OpenInfer Engine claims 2-3x faster inference than Llama.cpp and Ollama on distilled DeepSeek models and works as a drop-in replacement for existing endpoints. The company raised an $8M seed round in February 2025 and has since shipped Jean, a private email-native agentic AI system, and orchestration layers for heterogeneous compute.

- **Founded:** 2024
- **Headquarters:** San Mateo, California, United States
- **Founders:** Behnam Bastani (Co-Founder & CEO), Reza Nourai (Co-Founder & CTO)
- **Team size:** ~16 employees
- **Products:** OpenInfer Engine, Jean (AskJean.ai), Loom / Weave, OpenInfer API
- **Notable:** Raised an over-subscribed $8M seed round in February 2025 led by Cota Capital and Essence VC., Attracted angel backing from Google DeepMind Chief Scientist Jeff Dean, Oculus co-founder Brendan Iribe, and Microsoft CPO Aparna Chennapragada., Shipped the first preview build of the OpenInfer Engine with claimed 2-3x speedups over Ollama and Llama.cpp on distilled DeepSeek models.

## Products & services

- **OpenInfer Engine** — An AI agent inference engine that automatically disaggregates AI workloads and routes them to optimal hardware. Works as a drop-in replacement for existing endpoints and supports LangChain, Ollama and vLLM without code changes. Claims 2-3x faster inference than Llama.cpp and Ollama on distilled DeepSeek models.
- **Jean (AskJean.ai)** — A private, email-native agentic AI system that runs entirely on your own infrastructure with no cloud costs and no data exposure. Built on persistent contextual memory - a structured understanding of state, history and relationships. AskJean.ai is a public interface to try it.
- **Loom / Weave** — An orchestration and routing layer for enterprises running large models across fragmented, distributed compute. Weave continuously routes and schedules requests across heterogeneous compute in a closed loop, improving as more inference runs through it, with SLA-aware routing that matches each workload to the right compute topology.
- **OpenInfer API** — A zero-rewrite inference API that integrates into existing stacks as a drop-in replacement for existing inference endpoints.

## Achievements

- Raised an over-subscribed $8M seed round in February 2025 led by Cota Capital and Essence VC.
- Attracted angel backing from Google DeepMind Chief Scientist Jeff Dean, Oculus co-founder Brendan Iribe, and Microsoft CPO Aparna Chennapragada.
- Shipped the first preview build of the OpenInfer Engine with claimed 2-3x speedups over Ollama and Llama.cpp on distilled DeepSeek models.
- Launched Jean, a sovereign agentic AI system that runs entirely on user hardware.
- Announced partnerships with Intel and Microsoft in October 2025.
- Ran Llama 4 Scout locally in a client-side inference demo (April 2025).

## Latest updates

- **2026-06** — Published 'CPU: The Processor You Can't Route Around' and hosted a build-day hackathon; continued work on observability for heterogeneous inference.
- **2026-06** — Argued 'NVIDIA Dynamo Proved Inference Needs a Control Plane' - positioning Weave/Loom as the routing control plane for agentic workloads.
- **2026-04** — Launched Jean, a sovereign agentic AI system; hired a revenue chief; and published analysis tying agentic infrastructure inefficiency to Anthropic's Claude usage restrictions.
- **2025-10** — Announced partnerships with Intel (Partner Alliance) and Microsoft (Pegasus Program).
- **2025-02** — Closed $8M seed round; featured in VentureBeat.

## Links

- Website: https://openinfer.io
- LinkedIn: https://www.linkedin.com/company/openinfer

---

Profile page: https://yespress.io/openinfer
Published by YesPress — https://yespress.io
Last updated: 2026-07-03
