Breaking open the black boxAnthropic co-founder Chris Olah maps the hidden circuits of AIFrom DeepDream to mechanistic interpretabilityBreaking open the black boxAnthropic co-founder Chris Olah maps the hidden circuits of AIFrom DeepDream to mechanistic interpretability

Profile / Artificial Intelligence

Chris Olah Wants to Read the Mind We Built

For more than a decade, the Anthropic co-founder has pursued one stubborn question: if neural networks can surprise their makers, can we learn to see what they are actually doing?

The problem arrived disguised as a miracle. A neural network could recognize a face, translate a sentence, or tell a terrier from a teapot - yet nobody had written the rules by which it did so. The engineers had built the training process. The machine had built the method. Chris Olah could not get over this breach of ordinary programming etiquette.

Most software is a bureaucracy of instructions: every procedure has an author, every branch a reason. A trained network is more like a city seen from an airplane at night. You know the roads exist because the lights move, but the street plan is hidden in the glow. Olah made that opacity the center of his working life. His aim, as he puts it with almost suspicious neatness, is to “understand things clearly and explain them well.”

It is a compact credo for an untidy career. He left university without a degree, worked on open-source 3D-printing tools, entered machine learning through independent study and an exceptionally long internship, and became one of the people who gave a young field its pictures and its vocabulary. Today he is a co-founder of Anthropic and leads the company’s interpretability research. The title sounds administrative. The work is closer to natural history conducted on an organism made of numbers.

2012Thiel Fellowship
$100,000 for independent work
2 yearsThe Google Brain internship that kept going
2021Co-founded Anthropic and its interpretability program

A résumé made of evidence

Olah grew up in Toronto and graduated from the Abelard School in 2010 as an AP National Scholar. He studied mathematics at the University of Toronto, but university soon collided with a crisis at HackLab, the community technology space he had joined in high school. An acquaintance was arrested after police mistook materials in a hobby electronics and chemistry setup for explosives. Olah began attending hearings, taking notes, and helping the man’s family. He stepped away from university to follow the trial. The acquaintance was acquitted.

Olah has resisted turning this into a tidy tale of youthful heroism. He later said his help may not have altered the outcome much. The clearest contribution was practical: he transcribed an interrogation and synchronized the text to video, saving the legal team time. The episode nevertheless reveals a durable instinct. When a confident surface account may be wrong, inspect the record underneath it.

In 2012, a Thiel Fellowship gave him $100,000 and room to wander. His first destination was not artificial intelligence but 3D printing. He imagined open-source design software helping small machines reproduce useful objects. There was a printable vacuum cleaner and work on low-cost microscopes. Over dinner, Holden Karnofsky challenged the premise. Olah was irritated, then eventually grateful. The printers could reproduce many parts, but the few parts they could not make were precisely the difficult ones. Progress had acquired a flattering optical illusion.

“Your guesses about what’s going on inside a neural network will often be wrong.”Chris Olah, on the discipline of interpretability

Machine learning supplied a better obsession. Michael Nielsen, the physicist and science writer, was developing a book about neural networks in Toronto and convened a seminar to test its ideas. Olah joined just as deep learning was shaking off a long winter. He worked with Nielsen, studied on his own, and began emailing laboratories. One message to Yoshua Bengio took roughly a week to compose. It worked. So did a talk at Google, after which Jeff Dean offered him an internship at Google Brain.

The internship ran for about two years, a charmingly literal answer to the question of whether credentials or work matter more. Olah had no undergraduate diploma, but he had results, explanations, and an ability to turn abstraction into something another person could inspect. He also had timing: Google Brain was still a small group, and very few researchers were treating the hidden life of neural networks as a primary object of study.

Chris Olah discussing mechanistic interpretability with Lex Fridman
A black box meets two microphones: Olah explains mechanistic interpretability in a 2024 conversation with Lex Fridman. Image: Lex Clips / YouTube.

First came the strange pictures

In 2015, Alex Mordvintsev showed colleagues an early experiment that amplified patterns a vision network had learned. Olah abandoned his other projects and joined in. DeepDream turned ordinary photographs into feverish brocades of eyes, towers, feathers, and dog faces. The internet enjoyed the hallucinations. For Olah, they were a peephole. A network’s internal concepts, normally buried in matrices, could be coaxed into view.

The psychedelic phase matured into a visual science. Feature visualization asked what made a neuron respond. Activation atlases arranged a network’s learned concepts into explorable maps - something like a machine-made alphabet for images. Distill, the journal Olah co-founded, treated interactive diagrams as part of the argument rather than garnish applied after the difficult thinking was done. Its best articles let readers manipulate an idea until it clicked.

This was not communication bolted onto research. It was the method. Compress a network into one score and nearly all of its interesting structure vanishes. Preserve the detail in a good interface and patterns become available to the eye, then to experiment. Olah’s pictures were beautiful, but beauty was serving cross-examination.

At OpenAI, where he led interpretability research, the unit of attention shifted from neurons to circuits: groups of features carrying out recognizable computations together. The ambition hardened. It was no longer enough to say that a region of a network seemed associated with curves or animal heads. The task was to reverse-engineer an algorithm learned by training, tracing causes through the model much as one might trace instructions through compiled code.

The science before the textbook

Olah has called interpretability “pre-paradigmatic.” The field still argues over what understanding means, which questions matter, and what would count as a satisfactory answer. One camp may ask whether an explanation helps a person predict a model. Another may want a metric. Olah’s preference is more severe: treat interpretability as an empirical science whose claims can be shown true or false.

The distinction matters because a convincing story about a model can be completely wrong. Researchers are tempted to decide what they expect to find, probe for it, and congratulate the network for cooperating. Olah prefers methods that leave room for the rude fact. In one vision model, geography occupied far more of the last layer than he would have predicted. The surprise was the lesson. A tool for finding only expected problems is a lantern carefully aimed away from the dark.

DeepDream makes a vision network’s motifs visible.
Activation atlases organize learned visual concepts into landscapes.
The Circuits thread pursues reverse-engineered neural algorithms.
Anthropic maps millions of features in Claude 3 Sonnet.
Attribution graphs trace multi-step internal computations.
Natural-language autoencoders translate hidden activations into text.

Anthropic gave the program larger objects. In 2024, its researchers extracted millions of features from Claude 3 Sonnet. One responded to the Golden Gate Bridge. When they amplified it, Claude began inserting the bridge into answers where it had no earthly business appearing. The comic effect carried a serious point: the team had not merely found a correlation. It could intervene on an internal feature and change behavior.

By 2025, the group was tracing attribution graphs through language models. The resulting diagrams suggested internal steps behind planning, rhyme, arithmetic, and multilingual processing. They also showed the scale of the unfinished work. Even short prompts contain more computation than current methods capture, and interpreting the visible fraction can take hours. The microscope exists; the slide is the size of a country.

“If we could really understand these systems, we might be able to say when they are actually safe - or merely appear safe.”Olah on the long-term promise of model transparency

The model is not the only black box

In May 2026, Olah appeared at the Vatican for the presentation of Pope Leo XIV’s encyclical on artificial intelligence. An atheist AI founder addressing a papal gathering sounds like the setup to a niche joke, but his argument was sober. Frontier laboratories operate under commercial pressure, geopolitical pressure, pride, and ambition. Good intentions do not repeal incentives. The people building advanced systems need critics outside those incentives - religious communities, civil society, scholars, governments - who can say hard things.

The speech expanded interpretability from a technical practice into a civic one. Look inside the model. Look inside the institution. Ask not only what a system can do but how it does it; not only what a company promises but what pressures shape its choices. Olah spoke about the distribution of AI’s gains, the possibility of large-scale labor displacement, and questions of human flourishing that computer scientists are not qualified to settle alone.

There is an appealing consistency here. The young researcher who trusted diagrams more than incantations now asks his own industry to expose itself to inspection. Still, he is not selling certainty. His public language is full of caveats, partial maps, and work remaining. Interpretability has produced real mechanisms, but not a complete reading of a frontier model. It may contribute to safety; it has not furnished a safety certificate.

Away from the grand stakes, a lighter Olah occasionally appears. He has said an indulgent alternate career might involve teaching mathematics to young children - perhaps group theory as a board game, or knot theory as a guessing game. He invented the “micromarriage,” a playful unit for a one-in-a-million chance of meeting a future spouse, partly to coax himself toward social gatherings. The jokes work the way his diagrams do: give a slippery idea a form, then see whether the form helps you think.

That may be his most distinctive contribution. Plenty of people warn that artificial intelligence is opaque. Olah has spent his career making opacity specific. A feature can be isolated. A circuit can be traced. A hypothesis can fail. The mystery shrinks from a mood into a list of experiments.

The machines will not become simple because we have learned to draw them. But they may become legible enough to argue with. In a field crowded with forecasts, manifestos, and polished demonstrations, Chris Olah remains attached to a quieter proposition: before deciding what the machine is, first learn how to look.