01 $80M seed + Series A02 Sailboxes now generally available03 TPU v6e and Gaudi 3 research published01 $80M seed + Series A02 Sailboxes now generally available03 TPU v6e and Gaudi 3 research published

Company / 01 AI infrastructure

The AI Company That Learned to Wait

Sail Research has built an AI infrastructure business around a strange advantage: machines that can wait. Its bet is that the next useful agent will need more time, more tokens and a cheaper clock.

The agent was alive for 110 hours. It was working for only 64. That distinction, usually a footnote in a cloud bill, is close to the entire Sail Research argument. In a published demonstration, four of its cloud computers spent roughly 110 cumulative hours building a Redis-compatible program in Rust. The computers were awake for about 64 CPU hours; the rest was idle time. Sail says the workload cost $3.65 on its platform. A competing service, under Sail's comparison, would have cost about $9.30.

The short version
  • Sail sells model inference, persistent cloud computers called Sailboxes, and agent telemetry called Voyages.
  • Its preferred customer builds AI agents that research, code or run experiments for hours rather than answer a person in seconds.
  • The economic trick is to schedule patient model calls efficiently and let idle computers sleep.
  • The trade is occasional delay. A live voice assistant would have little use for patience.

The numbers in that demo belong to Sail, and the comparison is workload-specific. Yet the gap it identifies is real enough to build a company around: an autonomous program can spend much of its life waiting for a model response, a network request, another agent or a person. Ordinary infrastructure tends to measure elapsed time. Sail would rather measure work.

The customer who can afford to be patient

For years, the most visible AI product was a box on a screen. A human typed; a model answered; delay felt like failure. Inference systems were tuned accordingly. Neil Movva, Sail's co-founder and chief executive, and Samir Menon, its co-founder and chief technology officer, saw a different customer arriving: a software agent with a long assignment and no impatient person staring at a spinner.

Movva had worked on performance at NVIDIA, then at Apple and inference provider Together AI. Menon built systems at Apple. The two met as Stanford freshmen and reunited in late 2025. Movva told Fortune that the architecture at Together had been built for interactive applications. Their conclusion was to start over for background work. In Menon's words, the old goal was minimizing latency on one request; an agent's problem is sustaining throughput across many calls over hours.

The Sail Research team gathered around a chessboard in summer 2026
Sail's team, summer 2026. The chessboard is easier to spot than the unicycle, which is saying something.

Sail's model API makes the trade explicit. Developers can select a completion window for each request: default for a quick reply, balanced for background work, or flex for tasks that can wait. Sail's site advertises larger discounts as the window widens. This is less magic than scheduling. If a job does not have to jump to the front of the queue, the provider can pack hardware more efficiently and pass along some savings.

“We only care about efficiency.”Neil Movva, speaking to Fortune

A cloud computer with a snooze button

Inference is only half the bill. An agent that edits code, runs tests or conducts research needs somewhere to live between model calls. Sailboxes are persistent cloud computers for that work. They keep files and process state, can pause and resume, and can sleep automatically while waiting on Sail inference. Sail charges for active resources. Its published pricing also distinguishes CPU, memory and disk, so the exact bill depends on what a workload uses.

There is an engineering bill behind that billing model. Sail says its VMs can migrate between physical hosts while preserving open connections. It uses a network proxy to keep those connections stable, moves memory pages as needed, and lets memory expand or shrink with demand. The company first tried copying the entire VM through object storage. That was too slow for large machines. Demand paging, which fetches memory as the new host actually needs it, was the fix. Sail even ran a Minecraft server while forcing migrations every two minutes. It is a mischievous demonstration of a sober point: if moving a server does not ruin a game, a background agent can probably survive the occasional move.

Sail admits the cost strategy has a price. A command issued during migration can experience a few seconds of extra latency. The company expects migrations only a few times a day at most. For research running overnight, that may be harmless. For a live voice conversation, it could be the whole experience. The choice of workload matters more than the headline discount.

The third product is a record of the work

Sail added Voyages because a long-running agent creates a second problem: finding out what it actually did. A model may read files, call tools, launch subagents and leave artifacts, with no tidy account of the chain. Voyages records the model calls and sandbox commands in an attributed trace. Developers can compare versions of an agent workflow and inspect where a run slowed, failed or changed behavior.

Sail Voyages interface showing an attributed trace of agents, model calls and commands
The Voyages trace. When an agent works while everyone sleeps, somebody ought to keep minutes.
01 / Think

Sail Inference serves open models and lets a request declare how long it can wait.

02 / Work

Sailboxes give agents durable computers that sleep through idle periods.

03 / Explain

Voyages ties the model calls and commands back to the run that made them.

That stack reveals where Sail sits in the market. It is an infrastructure supplier to teams building agents, rather than an agent that writes your report itself. Detail.dev uses Sail for deep code-review agents; its chief executive says the company has used trillions of tokens. Parallel, which provides web context for AI agents, is another public customer. Quadrillion Labs uses Sailboxes for parallel experiments in Qualia Cloud. Sail also names Jack & Jill among its users. These are customers with repeated, token-heavy work where a few extra seconds can be exchanged for a smaller bill.

An $80 million wager on the clock

Sail announced $80 million in combined seed and Series A funding in June 2026, with Sequoia leading the seed and Kleiner Perkins leading the Series A. A company release put the valuation at $450 million. It sells usage-based inference and sandbox compute, with a self-serve tier, a $250-a-month Pro plan and custom enterprise terms. Sail says it already serves trillions of tokens a week, though it has not disclosed revenue.

The money funds an expensive kind of thrift. Sail's engineers tune inference software for particular chips, publish work on Intel Gaudi 3 and Google's TPU v6e, and orchestrate a fleet so less compute sits unused. None of that is useful if developers cannot adopt it, which explains the familiar OpenAI- and Anthropic-compatible interfaces. The cleverness has to disappear behind an API.

A builder can borrow the central idea without buying Sail. Measure an agent's active compute separately from its elapsed time. Ask which model calls truly need to be fast. Put durable state somewhere a task can resume after sleeping. Keep a trace that tells a human what happened. Sail has turned those habits into one commercial stack. Its future depends on whether enough valuable AI work can afford to take its time.