Profile Anastasios Angelopoulos Arena turns human choices into a live measure of AI Stanford engineer, Berkeley statistician, startup CEO

People / Artificial Intelligence / The Measurement Issue

The Statistician Who Put AI to a Public Vote

Anastasios Angelopoulos spent years asking how unreliable models could produce trustworthy decisions. Then a bare-bones Berkeley experiment turned millions of ordinary users into the jury for the AI industry.

The first thing to understand about Anastasios Angelopoulos is that he does not come to artificial intelligence from the usual direction. He did not begin with a chatbot persona, a consumer app, or a manifesto about machines replacing people. He came through uncertainty. At Stanford he studied electrical engineering. At UC Berkeley he worked on theoretical statistics and machine learning, asking how a person could make dependable claims around models that remain, in some essential way, black boxes. The problem was mathematical. It was also quietly practical: powerful systems fail, and people still have to decide what to do with them.

That question followed him from doctoral work into a plain-looking website built inside Berkeley's Sky Computing Lab. The interface offered two anonymous chatbot answers to the same user-written prompt. Pick the better one. Then discover which models had been competing. No ceremonial benchmark suite. No panel of judges in a sealed room. Just a choice, repeated by enough people to become evidence.

The intellectual path had taken years. At Stanford, advisers Gordon Wetzstein and Stephen Boyd introduced him to problems where mathematical structure meets engineered systems. At Berkeley, Michael I. Jordan and Jitendra Malik guided a doctorate that moved across statistics, machine learning, and vision. Angelopoulos collected an NSF Graduate Research Fellowship, a Berkeley Fellowship, and the Leon O. Chua Department Award. His papers returned to a consistent concern: when a model can be wrong in ways its designer cannot fully anticipate, what kind of guarantee can still be honest? That concern gave his later company a vocabulary before it had a cap table.

The experiment was called Chatbot Arena. It began in 2023 alongside the lab's work on open models and soon developed a life beyond the paper that described it. People arrived to test models on questions they actually cared about. Labs began watching the rankings. Some supplied unreleased systems under code names, turning the site into a preview room where the audience did not know the brand until after it had voted.

The Arena loop has the economy of a playground argument and the discipline of a controlled comparison: hide identity, hold the prompt steady, ask for a judgment.

Act I / A statistician walks into an arena

A moving test for a moving target

Static benchmarks offer comfort. They freeze a set of questions, grade the answers, and produce a number that fits neatly into a table. But once a test becomes important, people optimize for it. The questions leak into training data. Model builders rehearse the form. A measure designed to reveal capability can slowly become a measure of familiarity with the exam.

Arena's user-written prompts change the geometry. Every visitor can bring a new language, profession, irritation, joke, coding problem, or creative request. The distribution keeps moving because culture keeps moving. Angelopoulos has argued that a room of experts inventing a thousand prompts could not reproduce the range of things people ask in the wild. The useful test is the stream.

01 / AskA user supplies an organic prompt.
02 / HideModel identities stay concealed.
03 / ChooseThe user votes on the outputs.
04 / MeasureStatistics turn comparisons into ranks and uncertainty.

The design also borrows from a part of Angelopoulos's past. Before research papers and company formation, he was a national champion debater and a member of the US National Debate Team. Years later, he helped construct an arena where arguments arrive from masked participants and the audience decides. It is an irresistible symmetry, although the serious work happens after the spectacle. Pairwise preferences must be aggregated. Confidence intervals matter. Ties should remain ties when the evidence cannot support a clean separation.

“Everybody gets a voice.”Anastasios Angelopoulos on Arena's Wikipedia-like character

That democratic instinct does not mean a vote answers every question. People may reward confident prose, pleasant formatting, or an answer that feels helpful while missing a factual flaw. Angelopoulos has been explicit that preference evaluation does not replace every other form of testing. The interesting work is decomposition: accuracy, style, task completion, speed, taste, and domain expertise may need different instruments. The single leaderboard is a doorway, not a theory of everything.

Act II / Research escapes the lab

When the scoreboard became a company

By May 2025, the service was averaging one million monthly users across more than 100 languages. Angelopoulos and co-founder Wei-Lin Chiang had recently finished their Berkeley doctorates, become postdoctoral fellows, and were also housemates. The origin story still had the scale of an academic collaboration. The operating reality no longer did. Serving models costs money. Running constant evaluations requires infrastructure, product work, moderation, and relationships with the labs whose releases can redirect the crowd overnight.

Arena announced a $100 million seed round in 2025, introduced a commercial evaluation service that September, and followed with a $150 million Series A in January 2026. The name shortened from LMArena to Arena as the platform expanded beyond language-model chat into coding, search, images, video, and agents. The new name also sharpened the mission: measure intelligence where people use it, not only where researchers can prepare a clean answer key.

The velocity is unusual for a tool born in a university lab, but the handoff was not an abandonment of academia. The company continued publishing methods, data, and open-source components while recruiting specialists in research, infrastructure, product, and enterprise work. Angelopoulos moved from postdoctoral research with Ion Stoica and a student-researcher role at Google DeepMind into a job where inference bills, hiring, and customer needs share the calendar with statistical methodology. His title changed faster than his subject. The work still asks how evidence should be collected, weighted, and explained.

$250MCumulative funding announced across the 2025 seed and 2026 Series A
10M+Monthly visitors reported by Arena in June 2026
82M+Human votes reported by Arena in June 2026

The conversion from public research to private company introduces an obvious tension. Model labs can be subjects of Arena's rankings, customers of its evaluation products, and investors in the business around it. A referee can promise good intentions; a durable institution needs rules that survive personalities. Angelopoulos's answer has been structural. Public models are evaluated under the same methodology. A company cannot pay to enter the public leaderboard, leave it, or alter a score. The ranking pipeline is open, and statistical uncertainty is visible.

The useful thing to steal

If trust is part of the product, encode it in constraints outsiders can inspect. A neutrality claim gets stronger when revenue cannot purchase the outcome everyone watches.

This is also a business insight. Arena's commercial value depends on the credibility of its public signal. If the scoreboard looks purchasable, the audience leaves, the data becomes less representative, and the private evaluation product loses its reference point. Openness is not decoration around the model. It is part of the economic machinery.

Two public snapshots of monthly audience, thirteen months apart. The later figure was reported by Arena alongside its June 2026 business update.

Act III / Beyond the leaderboard

The suspicion behind the optimism

Angelopoulos's academic record makes his operating choices easier to read. His work on conformal prediction studies how to wrap predictions with statistically valid measures of uncertainty. Prediction-powered inference asks how machine-learning predictions can help answer scientific questions without allowing model errors to quietly become conclusions. Different papers, same instinct: use the model, but do not confuse its fluency with a guarantee.

That instinct has traveled into Arena's ranking method. Scores come with uncertainty. Close models may be tied. New categories reveal that a model strong at one task may be ordinary at another. In interviews, Angelopoulos has pushed against the idea that intelligence will collapse onto one universal axis. Taste persists. Context matters. A coding assistant, a search system, and a long-running agent may need different definitions of success.

His public persona has room for play around the formalism. Arena uses battles, crossed swords, secret model names, and the suspense of a reveal. Angelopoulos plays and records music with friends; he sings and handles guitar, piano, and bass. That detail belongs here because preference is partly an act of listening. Two correct answers can have different rhythm. A useful evaluator must notice both measurable failure and human response without pretending they are identical.

“Reality is the only benchmark that actually matters.”Anastasios Angelopoulos, 2026

Reality, of course, is expensive to measure. An agent might work for hours, call tools, write files, backtrack, and produce an outcome whose quality cannot be captured by a thumbs-up alone. Arena's recent work moves toward causal, post-deployment evaluation: compare systems doing actual tasks, inspect objective completion, and collect intermediate signals along the way. The ambition is broader than selecting the friendliest paragraph. It is to learn what happens when AI touches work.

The aspiration also stays practical. Angelopoulos has spoken about AI reducing repetitive paperwork and making workers more productive. His language is closer to engineering than prophecy. Build tools. Observe failures. Improve the test. Keep the platform accessible enough that real people continue supplying the weird, fresh edge cases no committee could schedule.

There is a final paradox in the story. Angelopoulos spent years developing formal tools for reliability, then helped build a system powered by something unruly: public taste. Arena works by giving that unruliness a shape without sanding it into false precision. Its votes do not settle what intelligence is. They record what people wanted from a model in a particular encounter, then let statistics say how much confidence that pile of encounters deserves.

The company will be judged by whether it can preserve that humility as the numbers, products, and stakes grow. The early interface made one elegant request of its users: look at the work before you look at the name. For a field saturated with brands, benchmarks, and launch-day claims, that remains a sturdy rule.