ai-evaluation

(10)
Legend
Wei-Lin Chiang and the People Who Grade the Machines
Founder · Engineer · Scientist

Wei-Lin Chiang and the People Who Grade the Machines

Wei-Lin Chiang helped turn a Berkeley experiment in comparing chatbots into Arena, where public votes now shape how AI models are measured. The engineer behind its systems keeps returning to a deceptively simple question: what do people actually prefer?

wei-lin-chiang · arena-aiRead →
Legend
The Quiet Debater Who Built a Business Around Judgment
Founder · Executive · Engineer

The Quiet Debater Who Built a Business Around Judgment

A quiet debate student became the engineer behind Mercor’s effort to find the people who can teach AI what expertise looks like. Now, as co-CEO, Adarsh Hiremath is trying to make that judgment work at a remarkable scale.

adarsh-hiremath · mercorRead →
Legend
Percy Liang and the Case for Showing the Work
Founder · Scientist · Engineer

Percy Liang and the Case for Showing the Work

Percy Liang has spent years asking how artificial intelligence should be built, tested and opened to scrutiny. From Stanford's foundation-model center to Together AI and Marin, his answer keeps returning to the same practical demand: show the work.

percy-liang · together-aiRead →
Legend
James Zammit Is Teaching Voice AI to Survive Tuesday
Founder · Engineer · Executive

James Zammit Is Teaching Voice AI to Survive Tuesday

After two startups arrived early and one finance product made precision non-negotiable, the Roark co-founder found his problem in the gap between a polished demo and an ordinary production call.

james-zammit · roark-aiRead →
Legend
Michael Malyuk Is Building the Judgment Layer AI Forgot
Founder · Engineer · Executive

Michael Malyuk Is Building the Judgment Layer AI Forgot

The HumanSignal co-founder turned a shared frustration with bad training data into Label Studio. His next wager is that smarter AI will need better human judgment, not less of it.

michael-malyuk · human-signalRead →
Legend
The Statistician Who Put AI to a Public Vote
Founder · Executive · Scientist

The Statistician Who Put AI to a Public Vote

Anastasios Angelopoulos spent years asking how unreliable models could produce trustworthy decisions. Then a bare-bones Berkeley experiment turned millions of ordinary users into the jury for the AI industry.

anastasios-angelopoulos · arena-aiRead →
Legend
Mohak Sharma
Founder · Executive · Operator

Mohak Sharma

Mohak Sharma is the co-founder and CEO of HoneyHive, a New York-based AI observability and evaluation platform that helps companies ship reliable AI agents to production. Built OpenTelemetry-native, HoneyHive lets engineering teams trace, test, and monitor multi-step AI pipelines and answer the deceptively hard question of what their agents are doing and why. Sharma started the company in 2022 with his former Columbia roommate Dhruv Singh, and by 2025 had raised $7.4M in total funding, including a $5.5M seed led by Insight Partners, while serving Fortune 100 customers in insurance and banking.

mohak-sharma · honeyhiveRead →
Legend
Jason Lopatecki
Founder · Executive · Engineer

Jason Lopatecki

Jason Lopatecki is the co-founder and CEO of Arize AI, a leading AI observability and evaluation platform that has raised $131M including a $70M Series C in 2025. A serial entrepreneur and UC Berkeley EECS graduate, he previously co-founded TubeMogul and scaled it from a garage startup to a NASDAQ-listed public company before Adobe acquired it in 2016. At Arize, he is building the infrastructure layer that helps engineering teams test, evaluate, and troubleshoot AI models and LLM-powered agents in production - a market that has exploded with the rise of generative AI.

ai · machine-learningRead →
Legend
Pankaj Gupta
Founder · Executive · Engineer

Pankaj Gupta

Pankaj Gupta is a serial entrepreneur and seasoned tech executive who built Twitter's early recommendation engine (Who-to-Follow, MagicRecs), led Google Pay's engineering across India and globally, scaled Coinbase's India operations from zero, and co-founded Yupp — a crypto-incentivized AI model evaluation platform that raised $33M from a16z before shutting down in March 2026. A Stanford PhD and IIT Delhi alumnus, he has founded four startups, three of which were acquired.

ai · machine-learningRead →
Legend
Hamel Husain
Engineer · Founder · Advisor

Hamel Husain

Hamel Husain is a machine learning engineer with 25+ years of experience who built part of the foundation beneath GitHub Copilot - his CodeSearchNet project was early LLM research later used by OpenAI for code understanding. Today he runs Parlance Labs, consults with AI teams across 35+ products, co-authored O'Reilly's 'Evals for AI Engineers', and teaches thousands of engineers how to move beyond vibes and actually measure their AI systems.

machine-learning · llmRead →