An AI model's answer is easy to see. The hard part is everything that happened before: the material it learned from, the decisions made during training, the tests it passed and the tests it failed. Percy Liang has built a career in that harder territory. He is a Stanford computer science professor, the founding director of its Center for Research on Foundation Models, and a founder of Together AI. His latest project, Marin, invites outsiders to watch model-building while it is still in progress. In a field fond of unveiling finished products, Liang keeps pointing to the workshop.
That interest predates today's chatbot boom. His personal site says he likes simple things, deep understanding and useful systems. It is an admirably plain description of work that can become very complicated very quickly. A language model can write a tidy paragraph in seconds; explaining exactly why it behaves as it does takes considerably longer. Liang's projects tend to make that slower work possible for more people.
Before the foundation, the score sheets
There is a delightful double entry in the honors section of Liang's Stanford homepage. One line records a silver medal at the 2000 International Olympiad in Informatics and second place at the 2002 ACM International Collegiate Programming Contest world finals. Another lists piano prizes: a Phoenix young musicians competition, MIT's concerto competition and a San Francisco radio station's classical search. The site offers no theory linking Chopin to computer science. It simply records that, for a while, Liang competed in both worlds.
He earned a bachelor's degree and a master's degree at MIT, then a doctorate at UC Berkeley, advised by Michael I. Jordan and Dan Klein. After a postdoctoral stint at Google in 2012, he joined Stanford's faculty. His earlier research crossed machine learning and natural language understanding: how to make a computer map a human request to something it can act on, and how to tell whether the method will hold up outside a neat demonstration. He later received a Sloan fellowship, an NSF CAREER award, the IJCAI Computers and Thought Award and a Presidential Early Career Award for Scientists and Engineers.
Those honors sit comfortably on a biography, but the tools reveal more about his habits. Liang created CodaLab Worksheets to let researchers manage experiments while keeping a trail from raw data to final result. A paper could become something a colleague could inspect and run, rather than a conclusion detached from its working parts. It is a rather unglamorous idea. So is keeping good lab notes. Both become more valuable when a result begins to matter.

A name for a change he could feel
By the time GPT-3 arrived, Liang had spent years working on systems designed for particular language tasks. In a Stanford conversation with colleague Chris Potts, he recalled expecting to find the new model's limits quickly. Instead, its flexibility surprised him. In another interview, he said the result “blew my socks off.” A system trained to predict text could answer questions, translate, summarize and shift among tasks in ways that seemed to require a different account of what AI had become.
The name his Stanford group helped popularize was “foundation model”: a broadly trained system that can be adapted to many uses. Liang later explained that “large language model” did not convey the role these systems would play as a base for other products and models. The term was a diagnosis as much as a label. If many tools rest on the same foundation, an error or a blind spot in that foundation can travel widely. The convenient new architecture also concentrates power in whoever can build and examine it.
“I am drawn to simple things, want to understand things deeply, and like to build useful systems.”Percy Liang, on his Stanford homepage
In 2021, Liang founded Stanford's Center for Research on Foundation Models with colleagues across disciplines. The center's early work included a large report on opportunities and risks. The breadth mattered. Computing systems that can be adapted to many domains pose questions for more than computer scientists. Liang described wanting the center to draw on expertise throughout academia, from the technical fields to the social sciences and policy.
A recurring question, six stops
The trouble with one number
A good score can settle an argument only if everyone agrees on the exam. Language models make that bargain difficult. A model may be strong at one task and weak at another; speed, cost, robustness and other concerns can pull in different directions. Liang and the CRFM team developed HELM, the Holistic Evaluation of Language Models, to make comparisons broader and more legible. Its premise is that an evaluation should reveal tradeoffs rather than polish them away into a single trophy.
The work also made something social visible: benchmarks help decide what a field tries to improve. When the Stanford team wrote about HELM, it called for evaluations that were pluralistic and open to scrutiny. They released code and detailed results so others could examine the model outputs beneath the aggregate scores. This is Liang's characteristic move. Put the measuring instrument on the table. Let someone else ask whether it measures the right thing.
A related effort, the Foundation Model Transparency Index, turned its attention to developers. The first version, introduced in 2023 by a multidisciplinary team that included Liang, examined 100 aspects of disclosure. It asked about the making of a model, its behavior and its downstream use. In the initial scoring of ten major developers, even the top results left ample room for more information. The index changed an airy word - transparency - into a set of questions that companies could answer or leave unanswered.
The professor and the company
Together AI, founded in 2022, places Liang beside researchers and engineers working on infrastructure for AI models. The company lists him as a founder. His Stanford work has its own academic aims; Together's platform serves teams building and running AI systems. The connection is visible in the question both settings face: who can participate when modern models demand large amounts of computing power and specialized tooling?
In interviews, Liang has argued that access to a model through a remote interface is only one kind of access. Researchers who want to study training choices, adapt a system or reproduce a result need more. A company that wants to deploy a model must also weigh trust, cost and where its data goes. His role at Together puts him close to the practical constraints behind those statements. An open recipe is useful only if people can actually work with it.
That is why his 2023 TEDAI talk asked for a more participatory way to build AI. The question is larger than whether weights can be downloaded. It includes whose work goes into a model, whose values shape it and who receives credit. The talk's public stage suited a researcher who has repeatedly moved technical arguments into wider view without pretending that a short speech settles them.
An experiment with the door open
Marin makes his newest proposal concrete. Created at Stanford in 2024 with David Hall and launched publicly in 2025, it is an open lab for building foundation models. Each proposed, running or completed experiment gets a GitHub issue. The issue states a goal up front, a small preregistration. A proposed run can then be submitted as code and reviewed in public. The site promises to share successful and unsuccessful experiments, data sources, model weights and discussions.
Failure is the notable item in that list. Finished AI models usually arrive with a story of progress. A public experiment log can show the detours, disagreements and ideas that did not work. That record may help a newcomer avoid repeating an expensive mistake, or notice a possibility the original team missed. It is also an invitation to criticize the work while criticism can still change it.
As of 2026, Marin says it is training a 535-billion-parameter mixture-of-experts model with 23 billion active parameters. The scale is striking, but the stronger statement is procedural: the project invites people to follow the run. Liang's personal page puts the aim plainly. Together with teaching Stanford's Language Models from Scratch course, he wants more people to understand, shape and contribute to foundation-model development. A lecture teaches the method; an open lab lets people watch it being tested.
“My goal is to enable everyone to understand, shape, and contribute to foundation model development.”Percy Liang, on teaching and Marin
This is not an easy standard to meet. Modern AI research is expensive, and an open process can be untidy. Marin's own public record makes those difficulties part of the account rather than a footnote. That is consistent with the path from CodaLab to HELM to the transparency index. In each project, Liang has tried to preserve or expose a piece of the chain between a claim and the work behind it.
The boy who could play a concerto and solve a programming contest problem appears only briefly in the public record. The professor's public work is much easier to trace. He builds instruments for checking claims, then asks other people to use them. In an industry where the final answer can be delivered almost instantly, he keeps making room for the slower, more human question: how did we get here?