Training gyms for AI agents - high-fidelity simulations of digital work where machines learn the job by doing the reps.
You wouldn't let a pilot fly a plane after only reading books about flying. That single analogy, repeated by co-founder and CEO Tim Lupo, is the whole thesis of Deeptune. The New York company builds what it calls "training gyms" - high-fidelity simulations of digital work that recreate the day-to-day software of an accountant, a lawyer, a customer-support rep, or a DevOps engineer. Inside those simulated workspaces, AI agents practice multi-step tasks and improve through reinforcement learning before they are ever pointed at a real system.
The problem Deeptune set out to solve is not that models are unintelligent. It is that intelligence alone does not make an agent reliable. A model can be brilliant on a benchmark and still fumble a five-step task in Salesforce, misfile a spreadsheet, or lose the thread of a support ticket. What closes that gap is practice - and practice requires an environment that looks and behaves exactly like the job. Deeptune's bet was to build those environments, bundling the problems, datasets, and infrastructure that let an agent fail safely, over and over, until it stops failing.
By the company's account it built hundreds of these environments for leading AI labs, and says they contributed to recent advances in agents' "computer use" - the move beyond simple question-answering into multi-step workflows on real software. It is an unglamorous layer of the AI stack. It also turned out to be a valuable one.
Deeptune's loop is simple to describe and hard to build: clone the software, hand the agent a real task, grade the attempt, repeat.
Recreate the tools a role uses all day - Slack, Salesforce, Excel, ticketing and monitoring systems.
Bundle realistic multi-step problems and datasets drawn from the actual workflow.
The agent attempts the task end-to-end inside the sandbox, where mistakes are free.
Reinforcement learning rewards good outcomes, turning capability into reliability over many reps.
High-fidelity simulations of digital workplaces where agents practice multi-step tasks and learn via reinforcement learning.
100+ ready-made environments bundling problems, data, and infrastructure - integrable in a few lines of code.
Environments modeled on real jobs: accountants, support reps, lawyers, and software/DevOps engineers.
While much of the industry argued over model benchmarks, Deeptune positioned itself one layer down - in the training infrastructure that turns capability into dependable behavior. Its customers were not consumers but frontier research labs and model developers, the same organizations racing to make agents useful at real work.
That placement is what differentiates it. Deeptune did not compete to build the smartest model; it built the gym where models get good at specific jobs. The environments are high-fidelity by design - close enough to the real software that skills learned inside transfer outside. Competing approaches range from human-data and evaluation vendors to other environment builders such as Mechanize, Prime Intellect, Surge AI, Scale AI and in-house lab teams. Deeptune's edge was depth of simulation and a small, senior team that had shipped this kind of software before.
Deeptune operated as B2B infrastructure. It built and licensed custom reinforcement-learning environments and evaluation datasets to frontier AI labs, which used them to train and benchmark agent capabilities. After the Mercor acquisition, that capability was folded into Mercor's reinforcement-learning stack - Mercor supplies domain experts and task-scoring systems, Deeptune supplied the simulated software environments where agents practice.
Public voice of the company and author of its flight-simulator thesis; framed New York as a deliberate recruiting edge for frontier AI.
Co-founded Deeptune to build simulation infrastructure for AI agents, part of a small team drawn from Anthropic, Scale AI, Palantir and Glean.
The roughly 20-person, in-person team included engineers and operators from Anthropic, Scale AI, Palantir, Hebbia, Glean, Retool and Modal.
Tim Lupo and Lukas Schmit begin building high-fidelity simulation environments for AI agents.
March round with 776, Abstract Ventures, Inspired Capital and angels to expand engineering and operations.
Deeptune's RL gyms contribute to advances in AI agents' computer-use capabilities.
In July, the AI unicorn acquires Deeptune and relocates the team to its New York office.
Roughly four months after the Series A, AI unicorn Mercor - valued at about $10 billion - acquired Deeptune. The deal had a tell: Mercor founder Brendan Foody had angel-invested in Deeptune's Series A months earlier. He described that bet plainly.
The acquisition consolidated Mercor's reinforcement-learning stack and moved Deeptune's entire team into Mercor's New York office. Financial terms were not disclosed.
It builds "training gyms" - high-fidelity simulation environments that recreate real workplace software so AI agents can practice multi-step digital tasks and improve through reinforcement learning.
It was founded in 2025 by Tim Lupo (CEO) and Lukas Schmit, based in New York City, with a small, in-person team drawn from Anthropic, Scale AI, Palantir, Glean and others.
A $43M Series A led by Andreessen Horowitz in March 2026, with 776, Abstract Ventures, Inspired Capital and angels; total funding is about $49.1M.
Frontier AI research labs and model developers, which use Deeptune's environments and datasets to train and evaluate agent capabilities such as "computer use."
In July 2026, about four months after its Series A, Deeptune was acquired by AI unicorn Mercor; the team moved to Mercor's New York office and its environments joined Mercor's reinforcement-learning stack.