Latest Anthropic Labs tests products at the edge of Claude2026 Mike Krieger joins Ben Mann in Labs2020 Mann helps build GPT-3's training systemsLatest Anthropic Labs tests products at the edge of Claude2026 Mike Krieger joins Ben Mann in Labs2020 Mann helps build GPT-3's training systems

Person / Engineering / AI safety

Ben Mann Keeps Asking What Happens Next

He helped build GPT-3, then helped found Anthropic. His current workshop tests what Claude can do - and what happens when a promising demonstration meets ordinary people.

In 2010, a young engineer named Ben Mann and a small group that included his brother tried to make an Android phone read Chinese characters and translate them without calling home to a server. They had a textbook, an idea, and a machine in their pockets that was not quite up to the assignment. The project ran out of momentum. The phones would catch up later. Mann kept the habit of testing the edge of what a machine could do, then paying attention when the edge pushed back.

That early failure makes a better introduction than a list of famous employers. In Mann's account, the problem was plainly physical: the phone lacked enough computing power. No amount of enthusiasm could make its processor a few years newer. It is a useful memory for someone now working with AI systems whose apparent fluency can make limits less visible. Sometimes the only honest result of an experiment is that the technology is not ready. Sometimes the right response is to try again when it is.

Mann describes himself on his personal site as a “software engineer, tinkerer, aspiring mad scientist.” The ordering is quietly revealing. There is play in the title, but engineering gets the first word. His path through Google, AI research and Anthropic has repeatedly brought him to the same practical question: what works when people are actually allowed to use it?

A startup plan interrupted by a book

Mann studied at Columbia University's Fu Foundation School of Engineering and Applied Science from 2007 to 2011, graduating magna cum laude. He then spent six years at Google. On an earlier personal site, he said his work covered the whole stack, from user experience to infrastructure scaling, and included leading a team of seven engineers. The span matters. A product can have an elegant screen and still break in the machinery behind it; a robust system can also be nearly impossible to use. Mann learned to care about both ends.

He left Google with a familiar ambition: start a company. While meeting potential cofounders and building side projects, he read Nick Bostrom's Superintelligence. Mann's own short biography gives the pivot with unusual economy. The book led him to focus on AI safety; he went to the Machine Intelligence Research Institute and OpenAI, and eventually helped found Anthropic. The prospective startup does not appear in that sentence as a near miss. It appears as an abandoned route to a problem that seemed larger.

This was not a retreat from making things. At OpenAI, Mann became one of the authors of the 2020 GPT-3 paper, Language Models are Few-Shot Learners. He is second on its author list. The paper's contribution note gives the less glamorous, more useful detail: Mann helped implement the large-scale models, training infrastructure and model-parallel strategies; he also worked on pre-training experiments, training data and downstream tasks. The public breakthrough depended on a great deal of plumbing.

6 yearsEngineering at Google before his AI safety turn
2020Second author on the GPT-3 paper
2021Anthropic is founded

GPT-3's headline capability was striking: it could perform some language tasks from examples supplied in the prompt, without a separate training run for each task. To the reader, that looked like asking a question. To the builders, it involved data preparation, experiments, scale and the difficult art of making an enormous system train at all. Mann worked on both the model and the conditions that let anyone learn what it could do. That combination - capability under scrutiny - would follow him into his next job.

The company built around a question

Mann was among the former OpenAI staff who founded Anthropic in 2021. The group wanted to study increasingly capable AI while developing systems that could be steered and trusted. Such words can sound abstract until a model sits in front of a user. It must answer the request, admit what it does not know, avoid unwanted actions, and remain understandable enough for someone to decide whether to rely on it. Those demands are technical, behavioral and commercial at once.

Mann's personal mission statement is direct: “My mission is to ensure transformative AI has a positive impact on humanity's long term future.” On the same page, he said that as of 2023 he was leading product engineering at Anthropic, where the company was trying to create a “safety race.” It is an intriguing phrase. A race usually rewards getting there first. Adding safety to the contest suggests a harder measure: whether a capable system behaves well enough to deserve repeated use.

“My mission is to ensure transformative AI has a positive impact on humanity's long term future.”Ben Mann, personal site

The work also made him a translator between different temperaments. Researchers want to know what a model can do and why. Product teams need to know whether the interface makes those abilities useful. People using it want the thing they asked for, preferably without the assistant decorating the request with surprises. In a 2025 conversation, Mann joked about coding models that made extra edits, comparing the experience to being offered fries and a milkshake when all you wanted was the original order. The joke lands because unwanted helpfulness is a familiar failure mode. A well-trained agent needs judgment about when to stop.

His discussions in 2025 also showed how he thinks about the clock. In another 2025 interview, Mann proposed an “economic Turing test” for broadly capable AI: whether systems can perform a substantial share of economically valuable work. He argued that people often underestimate the speed of change and spoke about the employment disruption he expected. These were his forecasts, not settled facts. Their relevance to his work is the practical pressure they create. If systems are improving quickly, a product team has to test capabilities that seemed implausible just a release earlier.

Ben Mann, seated at center with a microphone, speaking at an Anthropic discussion in Tokyo
A workbench with an audience: Mann, center, joins an Anthropic discussion in Tokyo in 2025.

A laboratory for ideas that might fail

In January 2026, Anthropic announced an expansion of Labs, its group for experimental products. Mike Krieger, Instagram's cofounder and Anthropic's former chief product officer, joined to build alongside Mann. Anthropic described a routine of trying rough versions with early users, seeing what survives, then handing successful ideas to teams equipped to make them dependable. It sounds like a familiar product cycle. At the edge of fast-changing AI, however, the material itself shifts while the product is being built.

The company pointed to Claude Code, the Model Context Protocol, Skills, Claude in Chrome and Cowork as products or technologies to come out of that approach. Its January announcement said Claude Code had grown from a research preview to a billion-dollar product in six months, while MCP had reached 100 million monthly downloads. Those figures are company claims, and they describe products rather than Mann's individual output. They do show why a group dedicated to uncertain experiments suddenly matters far beyond a small workshop.

In a September 2026 interview, Mann described Labs as a group of about 20 people that reviews prototypes on roughly two-week cycles. The team expects many ideas to fail or be folded into another project; Mann put the success rate at roughly 20 to 30 percent. That is a revealing number. A lab that celebrates every concept will soon have too many products and too little attention. A lab that kills the wrong one too early may miss the next Claude Code. The skill lies in learning fast enough to tell the difference.

The setup fits an old pattern in Mann's work. In his earlier project Webtendo, a laptop became a local game console and phones acted as controllers. His concern was not simply whether the arrangement could be assembled. It was whether play felt immediate: latency had to be low enough that the technology disappeared during a game with friends. The test was social. If people were cheering, arguing and playing again, the system had done its job.

Anthropic Labs has a more consequential audience. A coding assistant can reach into a real codebase. A tool connector can expose information or trigger actions. A desktop agent may operate across software that was designed for a human hand on the controls. The difference between a dazzling demonstration and a trustworthy product lives in small decisions: which action needs permission, which uncertainty should be shown, which extra edit must be avoided. Mann's instinct to test the edge now sits beside a reason to draw careful boundaries around it.

The useful answer is often “try it”

Mann's advice to companies in a 2025 interview was to experiment with uses that do not yet seem viable, measure how often they work and revisit them after new model releases. He argued that tomorrow's models may turn today's failed trial into a routine workflow. It is the opposite of writing a permanent verdict after one disappointing test. His Android translator, built before the phone was ready, gives the argument a biographical footnote.

There is an obvious tension in encouraging experiments while warning about AI risk. Mann does not resolve it by pretending the risk has disappeared. He argues for finding out what systems can do, building safeguards alongside them, and letting real use reveal problems that polished demos hide. The position asks for speed and restraint in the same room. It may be an uncomfortable room, but it is where products get made.

His public persona leaves room for some humor. “Aspiring mad scientist” is a good title for a person whose work now includes models that write code and use tools, but it is the word “aspiring” that keeps the joke honest. The experiments need a grown-up with a checklist as much as they need imagination. Mann's career has moved from a phone that could not yet recognize a character to systems capable of tackling much broader tasks. The question he keeps asking has not become simpler: after the machine shows that it can, what happens when everyone else tries?