The easiest way to make a language model look clever is to ask it a question it has seen before. This is a problem for anyone trying to buy one. Exam questions leak into training data. Model makers publish scores from tests they run themselves. A dazzling result can tell you little about whether an agent will finish a financial model, fix a bug, or assign a diagnosis code. Vals AI was built around a deliberately impolite response: let somebody else write the exam.
- Vals independently tests models, agents and AI products on professional tasks, then publishes performance, cost and speed.
- Its proprietary test sets stay private so models have a harder time memorizing the answers.
- Its strongest proof point is specific: leading models did well on clinical notes but far worse on medical billing codes.
- Enterprises can use its comparisons to choose models; developers can build a custom coding exam from their own GitHub history.
Co-founders Rayan Krishnan and Langston Nashold began asking how to measure the new generation of AI while at Stanford. Nashold has said he planned to join Hudson River Trading while Krishnan was preparing for PhD applications. They changed course as large language models arrived and the means of testing them lagged behind. Pear backed the pair in its 2023 accelerator cohort; the company launched publicly in 2024 with 15 models tested across four domains, keeping two datasets private.
The premise sounds almost quaint. The companies selling AI have made vast investments in getting better at tests. Vals wants to make a business out of writing better tests. It sells independent evaluations to model developers and enterprises, while its public leaderboards let anyone inspect a model against practical tasks. Its site tracks accuracy, operating cost and latency. A buyer can discover that the model with the best score may be too expensive or slow for the job at hand.

One profession, two very different grades
Consider healthcare. A model can write a plausible clinical note from a conversation. It can even put the observations in the correct SOAP format. But the same model may stumble when asked to assign an ICD-10-CM diagnosis code, where a single distinction in a condition hierarchy can change reimbursement and audit exposure. Vals tested both jobs rather than treating “medical AI” as one talent.
Its February 2026 study found leading models around 85 to 88 percent accurate on MedScribe, its clinical documentation benchmark. On MedCode, a medical billing test with 2,755 samples, most leading models clustered around 49 to 53 percent; the top result was 55.92 percent. The roughly 32-point gap is the company’s argument in miniature. A fluent answer and a correct operational decision are different accomplishments.
The healthcare split
To make that comparison, Vals worked with Protege on evaluation-ready, de-identified data and had certified professional coders annotate the coding samples. It used 80 transcripts and 100 expert-developed rubrics for the note-writing task. This is expensive, fussy work. It also exposes an awkward fact for AI procurement: the failure that matters most often hides behind the task the demo performed beautifully.
“We build benchmarks that measure the ability of models to do the work of lawyers, bankers, engineers, and doctors.”Vals AI, company statement
The hidden answer key
Vals divides its proprietary exams into public validation material, a larger private validation set it can license for internal work, and a final private test set used for published scores. That last set stays hidden. On open benchmarks, Vals runs models itself under a common harness and settings to improve comparability. It reports error bars as well as accuracy, cost and time. The method is a response to a common failure: once the answer key is on the web, the model may have read it during training.
There is a price for secrecy. A buyer can examine sample tasks and method notes, but cannot independently inspect every final question. Nor can one score settle every deployment decision: a law firm’s documents, a bank’s risk tolerance and a hospital’s coding rules are not interchangeable. Vals’s own methodology says its error bars cover uncertainty in the benchmark; they do not capture every variation in prompts, settings or model randomness. A company using the leaderboard still needs to test its actual workflow.
Those are the company’s own figures. Its August 2026 round was led by Andreessen Horowitz, with 8VC, Pear VC, Bloomberg Beta, HRT Ventures and NextLadder Ventures participating. Vals said its customer base had doubled and its team tripled in six months. Its results have appeared in model cards from OpenAI, Anthropic, Google, Meta and xAI. These are signs that independent grading is becoming useful to the very labs being graded, even when the grade is uncomfortable.
An exam made from your own mistakes
Vals Smith turns a customer’s GitHub repository into a coding benchmark. It looks for merged pull requests that represent clear, testable engineering work, rebuilds each as an issue with hidden tests, and rejects brittle candidates. The test must fail before the original fix and pass after it. A coding agent then receives the problem and the repository before the fix; it passes only if the hidden tests succeed without breaking existing behavior.
This is a practical idea a software team can copy even without buying a platform: gather real completed issues, reserve their fixes and tests, then put candidate agents through the same tasks. Vals Smith offers 120 starting credits. The company has not published a general enterprise price list, so the known cost is the effort of preparing good tasks and the compute consumed by runs, with paid evaluation terms handled directly.
The product’s failure mode is equally concrete. A repository with few cleanly testable pull requests will yield a thin exam. An agent that needs a complex production environment may be judged more on the setup than on its coding ability. And a benchmark drawn from last year’s issues cannot predict every bug next year. That is why Vals rejects tasks with brittle infrastructure and shows per-task breakdowns rather than only a trophy score.
A new kind of referee
The market already has public model scoreboards, academic evaluators and platforms for teams to run their own tests. Vals sits between them: part independent publisher, part benchmark studio, part infrastructure vendor. Its legal AI report compared four assistants across seven tasks against a lawyer baseline. Its finance and tax benchmarks ask agents to complete analyst work. Its open-source Valkyrie system runs agent evaluations at scale. The company has also moved into public-benefit questions, cybersecurity and other frontier risks.
The most interesting tension is that the referee is a business. Model labs and enterprises can pay for evaluations, while the public is meant to trust the scores. Vals’s answer is methodological discipline: independent runs, held-out tests, disclosed categories of tasks, common settings, rubrics and uncertainty estimates. Whether that is sufficient will be decided by the quality of its future exams and the usefulness of its failures, not by the round of funding.
The founders’ original question remains the useful one for any reader choosing an AI tool: what exact work must this system do, and what evidence shows it can do it? A general model can ace an exam and still miss the diagnosis code. The person buying it will discover which score mattered only after the task is done.