
Wei-Lin Chiang helped turn a Berkeley experiment in comparing chatbots into Arena, where public votes now shape how AI models are measured. The engineer behind its systems keeps returning to a deceptively simple question: what do people actually prefer?

A quiet debate student became the engineer behind Mercor’s effort to find the people who can teach AI what expertise looks like. Now, as co-CEO, Adarsh Hiremath is trying to make that judgment work at a remarkable scale.

Percy Liang has spent years asking how artificial intelligence should be built, tested and opened to scrutiny. From Stanford's foundation-model center to Together AI and Marin, his answer keeps returning to the same practical demand: show the work.

After two startups arrived early and one finance product made precision non-negotiable, the Roark co-founder found his problem in the gap between a polished demo and an ordinary production call.

The HumanSignal co-founder turned a shared frustration with bad training data into Label Studio. His next wager is that smarter AI will need better human judgment, not less of it.

Anastasios Angelopoulos spent years asking how unreliable models could produce trustworthy decisions. Then a bare-bones Berkeley experiment turned millions of ordinary users into the jury for the AI industry.
Mohak Sharma is the co-founder and CEO of HoneyHive, a New York-based AI observability and evaluation platform that helps companies ship reliable AI agents to production. Built OpenTelemetry-native, HoneyHive lets engineering teams trace, test, and monitor multi-step AI pipelines and answer the deceptively hard question of what their agents are doing and why. Sharma started the company in 2022 with his former Columbia roommate Dhruv Singh, and by 2025 had raised $7.4M in total funding, including a $5.5M seed led by Insight Partners, while serving Fortune 100 customers in insurance and banking.
Jason Lopatecki is the co-founder and CEO of Arize AI, a leading AI observability and evaluation platform that has raised $131M including a $70M Series C in 2025. A serial entrepreneur and UC Berkeley EECS graduate, he previously co-founded TubeMogul and scaled it from a garage startup to a NASDAQ-listed public company before Adobe acquired it in 2016. At Arize, he is building the infrastructure layer that helps engineering teams test, evaluate, and troubleshoot AI models and LLM-powered agents in production - a market that has exploded with the rise of generative AI.

Pankaj Gupta is a serial entrepreneur and seasoned tech executive who built Twitter's early recommendation engine (Who-to-Follow, MagicRecs), led Google Pay's engineering across India and globally, scaled Coinbase's India operations from zero, and co-founded Yupp — a crypto-incentivized AI model evaluation platform that raised $33M from a16z before shutting down in March 2026. A Stanford PhD and IIT Delhi alumnus, he has founded four startups, three of which were acquired.

Hamel Husain is a machine learning engineer with 25+ years of experience who built part of the foundation beneath GitHub Copilot - his CodeSearchNet project was early LLM research later used by OpenAI for code understanding. Today he runs Parlance Labs, consults with AI teams across 35+ products, co-authored O'Reilly's 'Evals for AI Engineers', and teaches thousands of engineers how to move beyond vibes and actually measure their AI systems.