Osmosis is a San Francisco reinforcement learning platform (YC W25) that helps companies fine-tune open-source models to beat foundation models on cost, latency, and reliability. Founded by Kasey Zhang and Andy Lyu, it lets teams train task-specific AI agents with multi-turn tool use over the Model Context Protocol, then export or self-host the weights they own.
Parasail is a San Mateo-based AI inference cloud that lets developers run open-source and custom models without buying or committing to their own GPUs. Instead of owning chips, it orchestrates rented GPU capacity across roughly 40 data centers in 15 countries plus its own clusters, automatically routing each workload to the cheapest, fastest hardware. Founded in 2023 by Mike Henry (ex-Groq CPO, Mythic founder) and Tim Harris (Swift Navigation founder), the company processes on the order of 500 billion tokens a day and pitches a token-based economy - pay for output, not hardware.
Nikola Borisov is the co-founder and CEO of Deep Infra, a Palo Alto cloud company that runs open-source AI models as a low-cost, low-latency inference service. A former competitive programmer who scaled messaging backends to hundreds of millions of users at imo.im and helped build HalloApp's founding backend, he bet early that inference, not training, would become the real bottleneck of the AI era. Deep Infra now processes close to five trillion tokens a week and raised a $107M Series B in May 2026 backed by NVIDIA, 500 Global, Felicis and others.
Deep Infra is a Palo Alto-based AI inference cloud that lets developers run hundreds of open-source machine learning models - large language models, image and video generation, speech, and embeddings - through a simple, OpenAI-compatible, pay-per-use API. The company owns and operates its own GPU fleet across multiple US data centers, processing trillions of tokens per week for companies that want production AI without managing infrastructure.
Hyperbolic is an open-access AI cloud built for developers and researchers. It aggregates underutilized GPUs from data centers, mining farms and idle machines into an on-demand marketplace, and pairs that with serverless inference for state-of-the-art open-source models. The pitch is simple: fast, affordable access to compute and inference without sales calls or feature gates, at prices it says can run up to 75% below incumbents.