PROFILE UPDATE   Andi Peng announced her departure from humans& in July 2026◆THE STORY   A career built around the question of human feedback◆PROFILE UPDATE   Andi Peng announced her departure from humans& in July 2026◆THE STORY   A career built around the question of human feedback◆

People / Artificial intelligence

Andi Peng and the art of asking what people want

From a Yale classroom and the White House to robot learning and a frontier AI startup, Peng has kept returning to one difficult question: what does the person on the other side of the machine actually want?

Consider the humble mug. To a person, a red one and a white one are both mugs. To a robot trained on the wrong examples, the color can become the whole story. Andi Peng, working with collaborators at MIT, asked what would happen if the machine could show a user why it failed and then learn which details were beside the point. It is a modest scene for a technical paper: one object, one mistake, one person invited to correct it. It is also a useful way into Peng’s career. Across policy offices, research labs, and a startup, she has kept returning to the gap between a system’s answer and a person’s intention.

That question long predates the current appetite for AI companies. At Yale, Peng chose two majors that sound as though they belong in different buildings: cognitive science and global affairs. In a later interview, she called herself an “academic mutt.” The joke carried a serious point. One discipline asked how individuals decide; the other looked at states and institutions making choices under pressure. Peng wanted to understand decisions in the real world, where the facts are incomplete and the consequences do not fit neatly into a multiple-choice exam.

The classroom before the lab

There was an early lesson in scale at Yale. In high school, Peng had seen Harvard’s CS50 while visiting a cousin and remembered the class’s energy. When Yale introduced its own version, she became head teaching assistant. The course drew 375 students that fall, more than its organizers expected. A room that size teaches an instructor something a benchmark cannot: students arrive with different backgrounds, different confidence, and different ways of asking for help. The same explanation that lights up one face may leave another blank.

In 2017, Peng was named a Truman Scholar. At the time, she was working at the intersection of technology and public policy, studying AI while also taking part in Yale’s Grand Strategy program. A summer at NASA had put her on a team that published a full-system engineering model of the Hyperloop. The topics look scattered in a résumé. Seen as questions about how technical designs meet human institutions, they sit closer together. Her undergraduate thesis studied terrorism with machine-learning methods. She was already crossing the line between social problems and computational tools.

2017Named a Truman Scholar at Yale
2023Published work on user feedback for robot adaptation
$480mSeed funding announced by humans& in January 2026

After graduation, Peng went to Washington and worked on federal science policy, including at the White House Office of Science and Technology Policy and the National Institute of Standards and Technology. She had also worked as a security engineer at Facebook. In a 2018 interview, she described seeing both the rapid social effects of technology and the difficulty of changing large organizations. The experience sharpened a decision: to contribute meaningfully to arguments about AI governance, she wanted deeper technical expertise.

Microsoft Research offered a bridge. As an AI resident, she joined a group studying collaborative decision-making with Eric Horvitz, Ece Kamar, and Emre Kiciman. The phrase can sound abstract until one asks what happens when a person receives a recommendation from a machine. Does it improve the decision? Does the machine’s bias move into the person’s judgment? Work Peng co-authored on AI-assisted hiring examined those questions directly. The person at the keyboard was part of the system under study, with all the independence and susceptibility that entails.

“How do decisions get made in the real world, especially when one is faced with risk and uncertainty?”Andi Peng, 2018

A robot that can be corrected

At MIT’s Computer Science and Artificial Intelligence Laboratory, Peng pursued the problem in a more physical form. She studied with Jacob Andreas and Julie Shah, and her work asked how robots might learn useful concepts from human feedback. A robot needs a compact picture of the world: what is a mug, what counts as a safe route, what detail matters to this particular user? The failure mode is almost comic. A machine can be impressively capable and still decide that a striped pan or a strangely colored cup is an entirely new species of thing.

Peng and collaborators built a framework for a robot that had failed a task to offer counterfactual explanations. Perhaps it would have succeeded if one visual feature were different. A user can then say whether that feature matters. The system uses the answer to generate new training examples and adapt the robot. Peng’s shorthand was memorable: she wanted to show it one mug, not 30,000. The point was to spare people the work of exhaustively demonstrating every color and pattern in the kitchen. The person provides the abstraction; the machine does the multiplying.

That research was presented at the International Conference on Machine Learning in 2023. Other MIT work looked at language as a way to help machines filter an environment. In a factory or kitchen, a robot sees a great many objects and properties; an instruction such as “hand me the bowl” contains clues about which ones matter. Peng and co-authors explored using ordinary language to build better representations for planning. It was another version of the same bet: human knowledge is not just a final score to optimize against. It can guide what a model notices in the first place.

Andi Peng speaking in an MIT CSAIL video interview
Peng discusses human feedback and robot learning in an MIT CSAIL spotlight. Frame from the lab’s video interview.

The subject suited her route through policy and research. An MIT profile noted that, after working in Washington, Peng had concluded she needed to become a technical expert to speak credibly about governance. Her work in robotics made that expertise tangible. It also kept a practical audience in view. A user at home does not want to write a reward function, diagnose distribution shift, or collect thousands of demonstrations. They want the robot to understand the task they meant to assign.

From one user to many

In 2024, Peng took leave from MIT and joined Anthropic. Her personal site says she worked on reinforcement learning and post-training for Claude versions 3.5 through 4.5. The scale had changed. A household robot has a body and a limited set of tasks; a language model meets many kinds of users and requests. Yet the feedback problem travels with it. What behavior should be rewarded? Whose preferences are represented? How should a system adapt when the request is ambiguous or the user’s intent changes?

By 2025, Peng had joined Eric Zelikman, Yuchen He, Georges Harik, and Noah Goodman to found humans&. The company introduced itself publicly in January 2026 with a claim that AI should help people communicate and coordinate. Its seed round was reported at $480 million, at a valuation of $4.48 billion. The number attracted attention, as numbers of that size do. The idea underneath was more specific: tools trained for one person asking one question may be poorly equipped for the untidy work of a group with different priorities, unfinished decisions, and a history it needs to remember.

One question, four settings
YALEHow do people and institutions make decisions?
MSRHow does machine advice change human judgment?
MITHow can a user correct what a robot misunderstands?
HUMANS&How might AI help groups coordinate over time?

In an interview about humans&, Peng described a plan to develop the model and its interface together. That is a consequential detail. A system built for coordination cannot simply be a very clever text box; it has to fit the human habits of asking, remembering, disagreeing, and revising. The team floated possible applications in work and everyday group decisions, but had not released a public product when the company announced the funding. The ambition was public; its practical form was still being worked out.

Peng announced in July 2026 that she was leaving humans&. Her public note expressed pride in the team and said she would cheer them on. The departure makes for a short founding chapter, and it gives the career a less convenient shape than a startup launch story usually allows. It does not erase the intellectual line that led her there. From hiring decisions to robot mistakes to the coordination of groups, Peng has repeatedly looked for ways to let people inform the systems built around them.

The question after the answer

A researcher’s biography is often written as a sequence of institutions. Peng’s does pass through a remarkable number of them: Yale, Washington, Microsoft, MIT, Anthropic, humans&. There are smaller details that make the sequence human. She says she read many dead philosophers as an undergraduate. She remembers being captivated by the spectacle of CS50. Her website asks people who write to her to use the subject line “Your cat is dope.” These are not clues to a secret personality; they are reminders that the people who design interactions have their own habits of interaction.

In an MIT interview, Peng said she would like to run a research group and teach students, whether through a university or in collaboration with industry. That aspiration fits the work already visible. Teaching, like building a useful AI system, depends on discovering what the other person understands and which explanation will help. The conversation has to remain open long enough for correction. There is a quiet symmetry in a researcher who once helped teach a packed introductory course later asking how a machine might learn to ask a better question.

The mug remains on the table. It is red, or white, or printed with a cartoon beaver; the color may matter to its owner, or it may be a distraction. A machine cannot settle that from performance alone. It has to ask, listen, and revise. Peng’s work has moved from single objects to broad AI systems, but this little domestic disagreement is still a good place to locate its stakes. The future of an intelligent tool may depend on its willingness to be told, politely and specifically, that it has picked up the wrong cup.