The first clue was a locked door. In 2022, Anand Kannappan heard that the investment bank where his brother worked had blocked access to OpenAI. He and Rebecca Qian had been friends since studying computer science together at the University of Chicago; both had gone on to work at Meta. To them, ChatGPT looked like a new kind of tool. To the bank, it looked like a new kind of liability. The founders' account of that mismatch became the opening argument for Patronus AI: enthusiasm for a model is one thing; permission to put it to work is another.
The short version
- Patronus AI tests model answers, compares versions and traces the mistakes of AI agents.
- Its customers include AI product teams at Gamma, Nova AI, Etsy and Algomo, according to company case studies.
- Its latest bet is a simulated digital workplace in which agents can train on long tasks before touching live systems.
The company was founded in 2023 and first sold a practical promise: automate the laborious checking that enterprises were doing with spreadsheets, manual reviewers and improvised prompts. Its software scores answers for things such as factual grounding and safety, runs experiments on datasets, and helps engineers compare a change against the previous version. That sounds technical until an assistant invents a figure in a financial document or misreads a customer request. Then evaluation becomes a matter of whether the product can be trusted at all.
The 81 percent problem
Patronus made its point with FinanceBench, a benchmark of questions grounded in public financial documents. The full dataset contains 10,231 questions with answers and supporting evidence. In a published sample of 150 questions, a GPT-4 Turbo retrieval setup answered incorrectly or refused to answer 81 percent of the time. The number describes that setup on that sample, not all uses of the model. Even with that limit, it made for an arresting demonstration: attaching a search system to a powerful model did not make it a dependable financial analyst.
FinanceBench also showed how Patronus differs from a generic dashboard. The team did not merely count responses. It worked with financial experts to ask questions whose answers could be checked against filings and earnings material. That research became a calling card for a platform with specialist evaluators, datasets and an API. Engineers can send an AI output for judgment, run a batch of examples, compare models or prompts, and look at the cases behind a score. Its Lynx evaluator addresses hallucinations in retrieval-based answers; multimodal judging extends the exercise to image-to-text work.

A score is useful. A cause is better.
A customer case gives the abstraction some weight. Gamma, the presentation platform, had more than 10,000 pieces of user feedback. Its team wanted to know why generated decks disappointed users, but there were too many long, open-ended outputs to label by hand. With Patronus, it picked a single error class - missing content - and built four evaluators to inspect factual completeness, length, structure and instruction following. The case study says Gamma saved more than 1,000 hours of manual evaluation per month and benchmarked more than 15 models.
The revealing detail is the sequence. Gamma did not begin with a magical score for “good presentation.” It named the failure, built checks for it, compared those checks with human judgments, and distilled a much smaller ground-truth set from the noisy feedback. The tactic is portable: choose an error that costs users something, collect real examples, check the grader against people, and rerun the same test after each change. A beautiful dashboard cannot rescue a test that measures the wrong thing.
“Patronus helped our AI team find signals and patterns of error in our datasets.”Jon Noronha, co-founder of Gamma
Other customers have different definitions of failure. Algomo used Lynx as a later check in a multilingual customer-support workflow and reported that precision on its internal hallucination set rose from 0.375 to 0.69. Etsy has used Patronus's image-to-text evaluation for product captions. Those are customer and company accounts, each tied to its own workflow. The common problem is ordinary and expensive: an answer can be fluent, formatted and wrong.
The first wrong turn
Agents changed the unit of inspection. An agent can search, call tools, write code and revise a plan before producing a final answer. If the final answer is poor, the mistake may have happened twenty steps earlier. Patronus's Percival reads execution traces and looks for failures in reasoning, planning, tool use and execution. It also proposes fixes. In a product example, an agent interpreted “today” as a hard-coded date from 2023. The final summary looked tidy; the search underneath it was stale. That is a more interesting bug than a spelling error, and harder to catch by grading the last sentence.

Nova AI's SAP migration agents supplied a tougher setting: runs lasting 20 to 30 minutes, with many model calls and specialized tools. A human reviewer needed roughly an hour to inspect one trace, according to Patronus's case study. Nova reported that Percival cut that work to a minute, helped fix three agent failures in a week, and raised accuracy on one SAP tool dataset by 60 percent. These are reported results from a specific deployment, not a universal speed guarantee. Their practical lesson is clear enough: preserve the full trace, identify the earliest consequential wrong turn, and test a concrete repair.
Now the test has walls and doors
In June 2026, Patronus announced a $50 million Series B led by Greenfield Partners, bringing disclosed funding to $70 million. It unveiled a preview of Digital World Models: generated environments that respond to an agent's actions as software might. Think of a simulated banking workflow, code repository or customer-service system. An agent can take an action, see a response and continue. In principle, that gives builders more than a pass-fail grade. It gives the agent a place to practice.
Set the world
Describe a task, tools, rules and the result that counts as success.
Run the agent
Let it navigate the generated environment and record the route.
Inspect and repeat
Find failed actions, vary conditions and compare the next run.
This is a larger bet than selling an evaluator API. Static test sets are useful, but they can only contain situations someone thought to write down. Hand-built agent environments can be costly and rigid. Patronus argues that generated worlds can vary tasks and failures more freely, turning evaluation data into training experience. Its preview invites teams to define an environment, connect through an API and run a simulation. The public announcement describes research and plans to spend the new capital on compute, infrastructure and an expanded research organization. The claim that a simulated workflow improves performance in a real one will depend on how faithfully the world models the tools, rules and surprises of actual work.
Patronus sits between several markets: AI observability, testing software, model benchmarks and now agent-training infrastructure. A team could also use in-house reviewers, open-source evaluation code or rival platforms such as LangSmith, Arize Phoenix and Galileo. Patronus's distinctive route has been to publish benchmarks and specialist evaluators, then build commercial tools around the problems those tests expose. Its customer stories suggest where that route works best: a team has enough real examples to define failure, enough volume to make manual review painful, and engineers willing to change prompts, tools or models in response.
The bank that barred ChatGPT in the founders' origin story was making a crude choice, but not a foolish one. The cost of a wrong answer was visible; the means to measure it were not. Patronus first offered a ruler. Then it offered a way to find the broken step. Now it wants to build the room where the step can be taken again. Whether that room resembles the workplace closely enough is the next test. At least this company has made a habit of asking for one.