The problem arrived disguised as a miracle. A neural network could recognize a face, translate a sentence, or tell a terrier from a teapot - yet nobody had written the rules by which it did so. The engineers had built the training process. The machine had built the method. Chris Olah could not get over this breach of ordinary programming etiquette.
Most software is a bureaucracy of instructions: every procedure has an author, every branch a reason. A trained network is more like a city seen from an airplane at night. You know the roads exist because the lights move, but the street plan is hidden in the glow. Olah made that opacity the center of his working life. His aim, as he puts it with almost suspicious neatness, is to “understand things clearly and explain them well.”
It is a compact credo for an untidy career. He left university without a degree, worked on open-source 3D-printing tools, entered machine learning through independent study and an exceptionally long internship, and became one of the people who gave a young field its pictures and its vocabulary. Today he is a co-founder of Anthropic and leads the company’s interpretability research. The title sounds administrative. The work is closer to natural history conducted on an organism made of numbers.
$100,000 for independent work
A résumé made of evidence
Olah grew up in Toronto and graduated from the Abelard School in 2010 as an AP National Scholar. He studied mathematics at the University of Toronto, but university soon collided with a crisis at HackLab, the community technology space he had joined in high school. An acquaintance was arrested after police mistook materials in a hobby electronics and chemistry setup for explosives. Olah began attending hearings, taking notes, and helping the man’s family. He stepped away from university to follow the trial. The acquaintance was acquitted.
Olah has resisted turning this into a tidy tale of youthful heroism. He later said his help may not have altered the outcome much. The clearest contribution was practical: he transcribed an interrogation and synchronized the text to video, saving the legal team time. The episode nevertheless reveals a durable instinct. When a confident surface account may be wrong, inspect the record underneath it.
In 2012, a Thiel Fellowship gave him $100,000 and room to wander. His first destination was not artificial intelligence but 3D printing. He imagined open-source design software helping small machines reproduce useful objects. There was a printable vacuum cleaner and work on low-cost microscopes. Over dinner, Holden Karnofsky challenged the premise. Olah was irritated, then eventually grateful. The printers could reproduce many parts, but the few parts they could not make were precisely the difficult ones. Progress had acquired a flattering optical illusion.
“Your guesses about what’s going on inside a neural network will often be wrong.”Chris Olah, on the discipline of interpretability
Machine learning supplied a better obsession. Michael Nielsen, the physicist and science writer, was developing a book about neural networks in Toronto and convened a seminar to test its ideas. Olah joined just as deep learning was shaking off a long winter. He worked with Nielsen, studied on his own, and began emailing laboratories. One message to Yoshua Bengio took roughly a week to compose. It worked. So did a talk at Google, after which Jeff Dean offered him an internship at Google Brain.
The internship ran for about two years, a charmingly literal answer to the question of whether credentials or work matter more. Olah had no undergraduate diploma, but he had results, explanations, and an ability to turn abstraction into something another person could inspect. He also had timing: Google Brain was still a small group, and very few researchers were treating the hidden life of neural networks as a primary object of study.
First came the strange pictures
In 2015, Alex Mordvintsev showed colleagues an early experiment that amplified patterns a vision network had learned. Olah abandoned his other projects and joined in. DeepDream turned ordinary photographs into feverish brocades of eyes, towers, feathers, and dog faces. The internet enjoyed the hallucinations. For Olah, they were a peephole. A network’s internal concepts, normally buried in matrices, could be coaxed into view.
The psychedelic phase matured into a visual science. Feature visualization asked what made a neuron respond. Activation atlases arranged a network’s learned concepts into explorable maps - something like a machine-made alphabet for images. Distill, the journal Olah co-founded, treated interactive diagrams as part of the argument rather than garnish applied after the difficult thinking was done. Its best articles let readers manipulate an idea until it clicked.
The progression is cumulative. Better pictures lead to candidate features; interventions test those features; connected features form circuits; circuits may expose why a model behaves as it does.
This was not communication bolted onto research. It was the method. Compress a network into one score and nearly all of its interesting structure vanishes. Preserve the detail in a good interface and patterns become available to the eye, then to experiment. Olah’s pictures were beautiful, but beauty was serving cross-examination.
At OpenAI, where he led interpretability research, the unit of attention shifted from neurons to circuits: groups of features carrying out recognizable computations together. The ambition hardened. It was no longer enough to say that a region of a network seemed associated with curves or animal heads. The task was to reverse-engineer an algorithm learned by training, tracing causes through the model much as one might trace instructions through compiled code.
The science before the textbook
Olah has called interpretability “pre-paradigmatic.” The field still argues over what understanding means, which questions matter, and what would count as a satisfactory answer. One camp may ask whether an explanation helps a person predict a model. Another may want a metric. Olah’s preference is more severe: treat interpretability as an empirical science whose claims can be shown true or false.
The distinction matters because a convincing story about a model can be completely wrong. Researchers are tempted to decide what they expect to find, probe for it, and congratulate the network for cooperating. Olah prefers methods that leave room for the rude fact. In one vision model, geography occupied far more of the last layer than he would have predicted. The surprise was the lesson. A tool for finding only expected problems is a lantern carefully aimed away from the dark.
Anthropic gave the program larger objects. In 2024, its researchers extracted millions of features from Claude 3 Sonnet. One responded to the Golden Gate Bridge. When they amplified it, Claude began inserting the bridge into answers where it had no earthly business appearing. The comic effect carried a serious point: the team had not merely found a correlation. It could intervene on an internal feature and change behavior.
By 2025, the group was tracing attribution graphs through language models. The resulting diagrams suggested internal steps behind planning, rhyme, arithmetic, and multilingual processing. They also showed the scale of the unfinished work. Even short prompts contain more computation than current methods capture, and interpreting the visible fraction can take hours. The microscope exists; the slide is the size of a country.
“If we could really understand these systems, we might be able to say when they are actually safe - or merely appear safe.”Olah on the long-term promise of model transparency
The model is not the only black box
In May 2026, Olah appeared at the Vatican for the presentation of Pope Leo XIV’s encyclical on artificial intelligence. An atheist AI founder addressing a papal gathering sounds like the setup to a niche joke, but his argument was sober. Frontier laboratories operate under commercial pressure, geopolitical pressure, pride, and ambition. Good intentions do not repeal incentives. The people building advanced systems need critics outside those incentives - religious communities, civil society, scholars, governments - who can say hard things.
The speech expanded interpretability from a technical practice into a civic one. Look inside the model. Look inside the institution. Ask not only what a system can do but how it does it; not only what a company promises but what pressures shape its choices. Olah spoke about the distribution of AI’s gains, the possibility of large-scale labor displacement, and questions of human flourishing that computer scientists are not qualified to settle alone.
There is an appealing consistency here. The young researcher who trusted diagrams more than incantations now asks his own industry to expose itself to inspection. Still, he is not selling certainty. His public language is full of caveats, partial maps, and work remaining. Interpretability has produced real mechanisms, but not a complete reading of a frontier model. It may contribute to safety; it has not furnished a safety certificate.
Away from the grand stakes, a lighter Olah occasionally appears. He has said an indulgent alternate career might involve teaching mathematics to young children - perhaps group theory as a board game, or knot theory as a guessing game. He invented the “micromarriage,” a playful unit for a one-in-a-million chance of meeting a future spouse, partly to coax himself toward social gatherings. The jokes work the way his diagrams do: give a slippery idea a form, then see whether the form helps you think.
That may be his most distinctive contribution. Plenty of people warn that artificial intelligence is opaque. Olah has spent his career making opacity specific. A feature can be isolated. A circuit can be traced. A hypothesis can fail. The mystery shrinks from a mood into a list of experiments.
The machines will not become simple because we have learned to draw them. But they may become legible enough to argue with. In a field crowded with forecasts, manifestos, and polished demonstrations, Chris Olah remains attached to a quieter proposition: before deciding what the machine is, first learn how to look.