
Wei-Lin Chiang helped turn a Berkeley experiment in comparing chatbots into Arena, where public votes now shape how AI models are measured. The engineer behind its systems keeps returning to a deceptively simple question: what do people actually prefer?

A quiet debate student became the engineer behind Mercor’s effort to find the people who can teach AI what expertise looks like. Now, as co-CEO, Adarsh Hiremath is trying to make that judgment work at a remarkable scale.

Percy Liang has spent years asking how artificial intelligence should be built, tested and opened to scrutiny. From Stanford's foundation-model center to Together AI and Marin, his answer keeps returning to the same practical demand: show the work.
Patronus AI began by catching a chatbot's bad answers. Now it is building practice worlds where agents can make their mistakes before a customer has to live with them.
A failed payment can masquerade as a change of heart. Global App Testing sends people into the markets, devices and conversations where software quietly lets users down.
The British testing company puts real people on real devices to find the faults that derail digital shopping. Its proposition: more eyes on a release, fewer surprises at checkout.

After two startups arrived early and one finance product made precision non-negotiable, the Roark co-founder found his problem in the gap between a polished demo and an ordinary production call.

The HumanSignal co-founder turned a shared frustration with bad training data into Label Studio. His next wager is that smarter AI will need better human judgment, not less of it.

Anastasios Angelopoulos spent years asking how unreliable models could produce trustworthy decisions. Then a bare-bones Berkeley experiment turned millions of ordinary users into the jury for the AI industry.
Nitrode is a San Francisco AI research company building the evaluation, data, and post-training layer for game development. It started in Y Combinator's Winter 2025 batch as a Godot-based AI game engine that let anyone prototype a 3D game in a day, then pivoted toward frontier research: it runs GameEngineBench, a benchmark that scores how well AI models handle real game-engine coding tasks, and produces expert-labeled datasets and reinforcement-learning environments (including Unreal Engine 5) for AI labs and enterprises. Nitrode's bet is that game development will be handled by systems of specialized models working together, not one general-purpose model.
Haize Labs is a New York-based AI safety and reliability startup that automates the red-teaming, stress-testing, and evaluation of large language models. Founded in 2024 by a trio of Harvard-trained researchers, the company builds algorithms that hunt for the inputs that make AI models misbehave - jailbreaks, failure modes, and edge cases - so they can be fixed before real users find them. Its 'haizing suite' and multi-turn attack engine Cascade are used by frontier model labs including OpenAI and Anthropic, alongside enterprises like Deloitte and MongoDB.
HoneyHive is a New York-based AI observability and evaluation platform that helps engineering teams ship reliable AI agents. Built OpenTelemetry-native, it unifies distributed tracing, offline experiments, real-time monitoring, and human-in-the-loop annotation into a single continuous-improvement loop so teams can see not just what an agent output, but how it got there and where it went wrong. Founded in 2022 by Columbia roommates Mohak Sharma and Dhruv Singh, HoneyHive is used in production by enterprises including Commonwealth Bank, NVIDIA, MongoDB, and Pinecone, and raised $7.4M in total seed funding led by Insight Partners.
Mohak Sharma is the co-founder and CEO of HoneyHive, a New York-based AI observability and evaluation platform that helps companies ship reliable AI agents to production. Built OpenTelemetry-native, HoneyHive lets engineering teams trace, test, and monitor multi-step AI pipelines and answer the deceptively hard question of what their agents are doing and why. Sharma started the company in 2022 with his former Columbia roommate Dhruv Singh, and by 2025 had raised $7.4M in total funding, including a $5.5M seed led by Insight Partners, while serving Fortune 100 customers in insurance and banking.
Vellum is a New York-based AI development platform that helps companies build, evaluate, deploy, and monitor production-grade LLM applications and AI agents. Founded in early 2023 by three engineers who kept hitting the same wall - prototypes that demo well but break in production - Vellum gives teams a workflow builder, an evaluation suite, version-controlled deployments, and live monitoring so they can move AI from proof-of-concept to reliable production systems. The platform works with more than 150 companies including Drata, Redfin, Swisscom, and Headspace, and raised a $20M Series A led by Leaders Fund in July 2025.
Braintrust is the end-to-end developer platform for shipping AI products. It connects evaluations, observability, prompt iteration, and an agentic optimizer (Loop) into one workflow so engineering teams can measure, debug, and improve LLM applications in production. Teams at Notion, Stripe, Vercel, Airtable, Instacart, Zapier, Ramp, Dropbox, Cloudflare and BILL use it to ship AI that doesn't silently regress.
Jason Lopatecki is the co-founder and CEO of Arize AI, a leading AI observability and evaluation platform that has raised $131M including a $70M Series C in 2025. A serial entrepreneur and UC Berkeley EECS graduate, he previously co-founded TubeMogul and scaled it from a garage startup to a NASDAQ-listed public company before Adobe acquired it in 2016. At Arize, he is building the infrastructure layer that helps engineering teams test, evaluate, and troubleshoot AI models and LLM-powered agents in production - a market that has exploded with the rise of generative AI.
Snorkel AI is a Redwood City-based enterprise AI company spun out of the Stanford AI Lab in 2019. Its Data Development Platform lets enterprises programmatically label, curate, and evaluate the training and evaluation data that powers custom LLMs, RAG systems, and AI agents - turning expensive manual annotation into a software-engineering discipline.

Pankaj Gupta is a serial entrepreneur and seasoned tech executive who built Twitter's early recommendation engine (Who-to-Follow, MagicRecs), led Google Pay's engineering across India and globally, scaled Coinbase's India operations from zero, and co-founded Yupp — a crypto-incentivized AI model evaluation platform that raised $33M from a16z before shutting down in March 2026. A Stanford PhD and IIT Delhi alumnus, he has founded four startups, three of which were acquired.

Hamel Husain is a machine learning engineer with 25+ years of experience who built part of the foundation beneath GitHub Copilot - his CodeSearchNet project was early LLM research later used by OpenAI for code understanding. Today he runs Parlance Labs, consults with AI teams across 35+ products, co-authored O'Reilly's 'Evals for AI Engineers', and teaches thousands of engineers how to move beyond vibes and actually measure their AI systems.