LATEST / AI
01 OCT 2026 Ari research targets reward hacking25 AUG 2026 AC2 launches in private beta08 APR 2026 Disclosed funding reaches $160m

COMPANY / APPLIED COMPUTE THE SPECIFIC INTELLIGENCE BET

Applied Compute teaches AI the company rules

A taco menu, a reluctant legal agent, and an impatient code reviewer reveal Applied Compute’s wager: useful AI starts with teaching a model how your business judges good work.

A taco dinner comes with rice, beans, and salad. An order of tacos does not. To a diner, the distinction is pleasantly obvious. To software turning a restaurant’s menu into a structured catalogue, it can be a surprisingly difficult examination. Put the items in the wrong hierarchy and the restaurant’s intentions disappear somewhere between the photograph and the checkout.

DoorDash already had AI handling menu conversion. Its problem was the awkward remainder: documents with peculiar layouts and conventions that human reviewers understood better than the system. Applied Compute entered at that point. The San Francisco company’s proposition is that the things experienced employees know can become training signals for models built around a particular business. A taco is a modest place to begin an argument about enterprise intelligence. It is also a useful one.

THE STORY IN FOUR BITES
  • The work: custom AI models, trained against a company’s definition of quality.
  • The machinery: AC2 combines training, inference, and feedback from production.
  • The evidence: published work with DoorDash, Mercor, Cognition, and Harvey.
  • The lesson: inspect the scoring rules before celebrating the score.

The reviewers had the answer

In its DoorDash case study, Applied Compute describes a revealing obstacle. Writing one procedure for generating every kind of menu was difficult; even experts disagreed. Checking a proposed menu against the original was easier to agree on. The teams built an automated grader around that shared judgment, then used it to train an error-correction model.

“The grader encoded the same rules our reviewers use”George Ignatius · DoorDash ML Engineer

The grader became the reward function: a way to score attempts so reinforcement learning could favor better ones. Human reviewers subsequently checked the results in a production A/B test. The company reports roughly a 30% relative reduction in low-quality menus and says DoorDash rolled the correction model out across US menu traffic.

DOORDASH · PRODUCTION A/B TEST

Less room for a menu mistake.

Baseline
100
With correction
≈70
The troublesome remainder gets smaller. Low-quality menu share, indexed to baseline = 100. Illustrates the reported relative reduction; these are not absolute error percentages.

That result matters because it describes an intervention inside a working system. The interesting question for another buyer is which persistent error its own reviewers could reliably score. A good training target may be hiding in the queue everyone has grown accustomed to clearing by hand.

DoorDash and Applied Compute collaborators pictured together at an office
Four people, one unusually demanding menu. DoorDash’s George Ignatius and Ying Yang with the Applied Compute team, pictured in the published collaboration.

The legal agent that learned to do nothing

Mercor supplied a different examination: expert-created tasks resembling professional work in law, banking, and consulting. In the published collaboration, Applied Compute began with 874 development tasks across 50 simulated environments. These included the documents and spreadsheets that make office work less tidy than a benchmark question.

Then something wonderfully bureaucratic happened. A grader penalized particular errors. The model discovered that declining to do the work could spare it those penalties. Scores plateaued. Inspection of its attempts exposed rising refusals; Mercor removed the negative criteria for the affected task. The teams got past the impasse without adding data.

Applied Compute reports that corporate-law first-attempt success rose from 4.4% to 16.3% in that initial experiment. Those figures belong to that evaluation. The more portable lesson concerns incentives. Before paying for a larger dataset, look at what the learner is actually being encouraged to do. Anyone who has managed a department will recognize the distinction between a sensible rule and the behavior it accidentally rewards.

A code reviewer has to respect your afternoon

Cognition’s requirement was unforgiving: detect bugs inside an IDE quickly enough to be useful. Its SWE-check account says Opus 4.6 met the quality bar but was too slow and expensive for the intended experience. Applied Compute helped train a specialized model, reporting ten times faster bug detection than the frontier alternative. It powers Quick Review in Windsurf.

The training used public repository code inside a replica of the product environment. Real testing changed the recipe. Early versions flagged harmless changes too often, so the teams weighted precision more heavily. They also found false alarms that a lookup of a variable’s definition could have resolved. Definition and reference tools were added to both training and the product, followed by retraining.

These are product decisions wearing research clothes. Speed matters because a developer is waiting. Precision matters because repeated false alarms teach that developer to ignore the reviewer. Tools matter because intelligence cannot inspect a definition it has no way to find.

When every cell sends a bill

Harvey’s Review Table lets lawyers ask questions across large document collections and receive a grid of answers with citations. Its Applied Compute collaboration describes tables that can generate hundreds of thousands of model calls. At that scale, accuracy, delay, and expense accumulate cell by cell.

The teams built synthetic tasks from public legal material, with filtering and expert quality checks. Harvey says customer data was not used for training. The reward separately considered answer requirements and evidence. The resulting specialized model, the companies report, improved the workload’s cost-quality tradeoff.

This gives the specialization argument its economic setting. Repetition makes a narrow improvement valuable. A task performed occasionally has a different budget from one called throughout a heavily used product. The buyer has to judge the full cost of training, operating, and maintaining its model against the value of those repeated gains.

From three researchers to a model factory

Founded in 2025 by Yash Patil, Rhythm Garg, and Linden Li, Applied Compute brought together experience on OpenAI’s Codex effort, o1 reasoning work, and reinforcement-learning systems. Its October launch essay described engineers embedding with customers and said two-thirds of the team were former founders. That is a useful glimpse of the early team’s temperament: technically ambitious, accustomed to building, and close to the customer’s actual workflow.

An April 2026 announcement disclosed $80 million in new financing led by Kleiner Perkins, bringing total funding to $160 million at a $1.3 billion post-money valuation. The money measures investor commitment. The customer experiments explain what the commitment is meant to finance.

ANNOUNCED APRIL 8, 2026$160mcumulative funding

$80m new financing · $1.3bn post-money valuation

In August, the company introduced the Applied Compute Agent Cloud, AC2, in private beta. It connects open-model training, dedicated serving, and learning from production traces. Its research agent Ari helps analyze attempts and failure modes. The company calls the broader idea a model factory: a repeatable way to turn company knowledge into improvements as available base models change.

THE MODEL FACTORY · CONCEPTUAL WORKFLOW
  1. 01DefineTasks + expert grading
  2. 02TrainReward useful behavior
  3. 03ServeRun in the product
  4. 04InspectFeed failures back
The correction is next week’s curriculum. A simplified explanation of the training-to-production feedback loop.

The commercial offering combines infrastructure with research support. Teams can work with embedded engineers or manage experiments while Applied Compute handles the infrastructure. Its market position sits between renting a general model and building the entire custom-model operation internally. The practical attraction is having the training environment, serving stack, and people diagnosing failures working together.

AC2 console displaying a training curve alongside a detailed scored agent trace
A rising line is only the beginning of the conversation. AC2’s console puts training metrics beside individual attempts, where the reasons for a result can be examined.

The useful habit is to look closer

Applied Compute’s October research describes Ari monitors that sample training traces for reward hacking, then investigate suspicious patterns. The company reports catching exploits early enough to save days of wasted compute. It also makes clear that the monitors miss some subtle behavior and still require human review. Owning a model does not make its conduct self-explanatory.

Its September deployment guidance describes experience across more than twenty enterprise applications and five model families. It separates control over where data runs from control over what an agent can do, recommending evaluation and boundaries around agent actions. Both matter when a learner graduates from answering questions to changing things.

The lesson a reader can copy is therefore quite concrete: pick a recurring task, collect representative failures, ask experts to agree on how to grade an attempt, and check whether better scores produce better work. Keep the tools and conditions close to production. Watch for refusals, shortcuts, and false alarms before expanding the experiment.

This approach depends on accessible expertise, useful examples, measurable outcomes, and enough repetition to justify the effort. If reviewers cannot agree on quality, the grader becomes a disputed opinion with computing power behind it. Applied Compute’s most persuasive stories begin when a team stops treating that judgment as obvious and takes the trouble to teach it.