Profile / Wei-Lin Chiang◆Berkeley lab to public leaderboard◆Two anonymous answers, one human vote◆Now: Arena co-founder and CTO◆Profile / Wei-Lin Chiang◆Berkeley lab to public leaderboard◆Two anonymous answers, one human vote◆Now: Arena co-founder and CTO◆

People / The measure of AI

Wei-Lin Chiang and the People Who Grade the Machines

Wei-Lin Chiang helped turn a Berkeley experiment in comparing chatbots into Arena, where public votes now shape how AI models are measured. The engineer behind its systems keeps returning to a deceptively simple question: what do people actually prefer?

A stranger opens a web page and types a question. Two chatbots answer, but their names are hidden. The stranger picks the response that works better. Only after the vote do the names appear. There is no lab coat in sight, no examination booklet, and usually no reason for the voter to know the statistical machinery humming underneath. Yet that small act has become part of how the AI industry learns what people think of its products.

Wei-Lin Chiang helped build the machinery. In 2023, he and fellow researchers at UC Berkeley launched Chatbot Arena as an open way to compare language models. The interface was spare enough to invite anyone in. The premise was ambitious enough to keep researchers busy: if the questions come from the public, and the judgments do too, the resulting score may reveal things a fixed test cannot. Three years later, the project is a company called Arena, and Chiang is its co-founder and chief technology officer.

A leaderboard is a peculiar kind of fame. Its rows are named for other people's machines. The person who keeps the contest running can almost disappear behind them. Chiang's career makes better sense if you look beneath the ranks, at the systems that make a comparison possible and the uncomfortable questions that arrive once a comparison starts to matter.

Before the arena, the plumbing

Chiang studied computer science at National Taiwan University, earning a bachelor's degree with a mathematics minor and then a master's degree. His master's thesis dealt with training large graph convolutional networks efficiently. It is the sort of work that puts scale and practicality in the foreground: what happens when an elegant method has to run on a large, unruly problem? His publication record from that period includes work on distributed optimization and Cluster-GCN, an approach to training large graph networks.

The workplaces on his personal site read like a tour of AI infrastructure. Internships took him to Microsoft in Redmond, Microsoft Research Asia in Beijing, Alibaba in Hangzhou, Google Research in Mountain View, and later Amazon in Seattle. The subjects varied, from distributed training to information extraction. The common concern was how to make computationally demanding ideas operate at useful scale. His 2018 paper with Chih-Yang Hsia and Chih-Jen Lin won a best paper award at the Asian Conference on Machine Learning. This was systems work with a stubborn eye for implementation.

At Berkeley, Chiang joined the Sky Computing Lab and worked with Ion Stoica. He contributed to SkyPilot, which helps run AI workloads across clouds, and to FastChat, the open-source framework that served models for Vicuna and Chatbot Arena. Vicuna, an open chatbot developed by the Berkeley group, sharpened the next problem. Once you have a model to show, how do you know how it compares with the alternatives? A published benchmark can provide one answer. A person trying to use the model might provide another.

The simple vote, step by step
01 / ASKA person writes a prompt
02 / HIDETwo unnamed models respond
03 / CHOOSEThe person picks one answer
04 / COUNTThe vote informs a ranking
The identities are revealed after the choice, limiting the influence of a model's brand on the initial vote.

A test written by its takers

The first Arena results appeared in May 2023. The researchers described anonymous, randomized battles and an Elo-style ranking, borrowing a familiar idea from competitive games. The academic paper Chiang co-authored the following year laid out the case for crowdsourced human preference and reported more than 240,000 votes by that point. These numbers were soon overtaken, but the central move remained unchanged. The people using the system wrote the questions. They could ask for help with code, a translation, a joke, a difficult explanation, or something nobody had thought to put on a standard exam.

That freedom is scientifically useful and scientifically awkward. A fixed test is tidy. You can repeat it, compare results and know exactly which questions were asked. An open test sees more of ordinary life, but its voters bring different tastes, languages and habits. A clever answer may be more persuasive than a correct one. A long answer may look better than a concise answer even when it says less. Chiang and his collaborators have studied these problems, from the statistical ranking of paired responses to the influence of style and the freshness of submitted prompts.

“We want people to test these models and express their opinions and preferences”Wei-Lin Chiang, 2025 interview

The quote is a useful description of the project, and also a warning about its limits. Preference is a signal about what a person liked in a particular comparison. It is not a universal verdict on every ability of a model. The ranking has to be read with an understanding of the prompts, the voters and the task. That is why the plain-looking vote rests on so much less visible research. A popular scoreboard can make measurement feel settled just when the underlying questions are becoming more interesting.

Wei-Lin Chiang with Anastasios Angelopoulos and Berkeley student collaborators around a restaurant table
From left: Evan Frick, Tianle Li, Aryan Vichare, Wei-Lin Chiang and Anastasios Angelopoulos, pictured together during the Berkeley chapter of the project. A leaderboard can begin with a ramen table.

A crowd, a company, a credibility test

Berkeley described Chiang and co-founder Anastasios Angelopoulos as housemates when Chatbot Arena was growing. It also pictured them with three student collaborators, Evan Frick, Tianle Li and Aryan Vichare, around a restaurant table. The photograph is an antidote to the idea that a consequential internet product arrives fully formed. There are bowls, chopsticks and five people who had been working on the same difficult problem. The site itself began as a bare text interface. No marble columns were required to invite the public to vote.

By spring 2025, the site was drawing roughly a million monthly users, according to Berkeley. LMArena then became a company, with Chiang as CTO and Angelopoulos as CEO. In May, it announced a $100 million seed round. The public platform would continue to be free, while the organization gained resources to expand its infrastructure and research. That move made an old tension more visible. A benchmark that influences model makers can also do business with them. The founders now have to make the ranking's methods and independence credible to people who use it and people who are ranked by it.

The argument over neutrality is not an abstract footnote. A shift of a few places on a public leaderboard can become a marketing line. A model maker may submit versions before a launch. A crowd can favor a confident tone, a particular language or a polished format. Researchers have raised questions about gaming and representativeness. Chiang's response has been to keep extending the measurement itself: study fresh prompts, separate style from substance, publish methods, and take the same comparative format into different sorts of tasks. None of that removes every judgment call. It does make those calls harder to hide.

10M+Monthly visitors
700M+Total conversations
82M+Total votes

Company figures reported June 2026. These are platform totals, not personal metrics.

When the question changes

Chat was only the opening round. Chiang's project list includes multimodal comparisons, in which users can submit images, and work on evaluating coding assistants, image generation, search and agents. The expansion follows the technology: once AI can inspect a picture or edit software, a text-only answer misses part of the performance. Each new format brings its own measure of usefulness. A pleasing image, a sound code edit and a completed multi-step task are not interchangeable, however tempting it may be to flatten them into one score.

He has also worked on a problem that haunts every public exam: when the questions are known, the students can practice them. Language models train on enormous amounts of text, so a fixed benchmark may cease to be a clean test. In research with collaborators, Chiang examined how even rephrased test samples complicate claims about contamination. Later work on Arena's user prompts found that most sampled prompts were fresh, a reason a live stream of questions can provide a different view from an old question bank. The answer is not to worship novelty; it is to notice when an apparently precise score rests on familiar material.

In January 2026, LMArena shortened its name to Arena. The name had been there from the beginning, in the idea of two models meeting in a public contest. The company said the new name reflected a platform broader than language models. By June, it reported more than 10 million monthly visitors and 82 million cumulative votes. Those figures describe scale. They also describe responsibility: at that size, the tiny interaction Chiang helped build is no longer merely an experiment. It is part of the public grammar for talking about AI progress.

The human at the end of the line

Chiang has spoken of wanting people to test models and express their preferences. It is a practical aspiration, less theatrical than predictions about machines surpassing everyone. The work begins with a user's question and ends with a user's judgment. Between those moments are servers, sampling choices, statistical estimates, interface details and plenty of opportunities for error. The engineer's job is to make that path short enough for anyone to use and rigorous enough to deserve attention.

There is a modest joke in the arrangement. Elaborate AI systems enter a contest, and the decisive gesture is a person clicking the answer they prefer. The joke lasts only until you ask what that click means. Chiang has spent years building the systems that let researchers ask that question carefully. In a field fond of announcing winners, his work lives in the more durable business of deciding how a winner is chosen.