Modal is a serverless cloud platform that lets developers run compute-intensive AI and data workloads - inference, training, fine-tuning, batch jobs, and sandboxed code - without managing infrastructure. Built from scratch in Rust with a custom container runtime, file system, and GPU memory snapshotting, it delivers sub-second cold starts and elastic autoscaling from zero to thousands of GPUs, billed by the second. Founded in 2021 by Erik Bernhardsson and Akshat Bubna, the New York-based company reached roughly $300M in annualized revenue and raised a $355M Series C in 2026 at a $4.65B valuation.
Erik Bernhardsson is the founder and CEO of Modal Labs, a serverless cloud platform that gives AI engineers on-demand GPU compute with sub-second container starts. A Swedish engineer trained in physics at KTH, he built the first version of Spotify's music recommendation system (including the groundwork behind Discover Weekly), open-sourced the widely used Luigi and Annoy libraries, then ran a 300-person engineering org at Better.com before starting Modal in 2021. He is also a prolific blogger whose essays on engineering, data teams, and startups reach a large technical audience. In 2026 Modal raised a $355M Series C at a $4.65B valuation.
Nikola Borisov is the co-founder and CEO of Deep Infra, a Palo Alto cloud company that runs open-source AI models as a low-cost, low-latency inference service. A former competitive programmer who scaled messaging backends to hundreds of millions of users at imo.im and helped build HalloApp's founding backend, he bet early that inference, not training, would become the real bottleneck of the AI era. Deep Infra now processes close to five trillion tokens a week and raised a $107M Series B in May 2026 backed by NVIDIA, 500 Global, Felicis and others.
Deep Infra is a Palo Alto-based AI inference cloud that lets developers run hundreds of open-source machine learning models - large language models, image and video generation, speech, and embeddings - through a simple, OpenAI-compatible, pay-per-use API. The company owns and operates its own GPU fleet across multiple US data centers, processing trillions of tokens per week for companies that want production AI without managing infrastructure.