# Chris Olah

> Chris Olah is a Canadian machine-learning researcher, Anthropic co-founder, and interpretability research lead who has spent more than a decade developing visual and mechanistic ways to understand what neural networks do internally. After leaving university without a degree, he pursued independent work, became a 2012 Thiel Fellow, joined Google Brain through an unusually long internship, led interpretability research at OpenAI, co-founded the visual research journal Distill, and helped establish mechanistic interpretability as a distinct empirical program.

- **Role:** Co-Founder and Interpretability Research Lead at Anthropic
- **Organizations:** Anthropic, OpenAI, Google Brain, Distill, Thiel Fellowship
- **Nationality:** Canadian
- **Education:** High school diploma; AP National Scholar, The Abelard School, Mathematics studies; left without a degree, University of Toronto
- **Known for:** Helped establish mechanistic interpretability as a distinct approach to reverse-engineering learned neural-network computations., Co-founded Anthropic in 2021 and leads its interpretability research., Co-founded Distill, which made interactive visualization and clear exposition central to communicating machine-learning research.

## Career timeline

- **2010** — Graduated from The Abelard School in Toronto as an AP National Scholar with six AP credits.
- **2012** — Received a $100,000 Thiel Fellowship to pursue independent work; initially explored open-source tools for 3D printing.
- **2013-2014** — Shifted into machine learning after participating in Michael Nielsen's Toronto seminar on neural networks and began visiting research groups.
- **2014** — Joined Google Brain as an intern after giving a talk at Google; the internship ultimately lasted about two years.
- **2015** — Contributed to Google Brain's DeepDream work and continued developing neural-network visualization techniques.
- **2016** — Co-authored Concrete Problems in AI Safety and joined OpenAI, later leading its interpretability research.
- **2017-2018** — Co-founded Distill, an online research journal centered on clear, interactive technical communication.
- **2019** — Co-developed activation atlases, a visual technique for exploring concepts represented inside image-classification networks.
- **2020** — Led the OpenAI Circuits thread, advancing the program that became known as mechanistic interpretability.
- **2021** — Co-founded Anthropic and began leading its interpretability research.
- **2024** — Anthropic's team mapped millions of interpretable features in Claude 3 Sonnet; TIME named Olah to the TIME100 AI list.
- **2025** — Co-authored work tracing internal circuits and attention computations in language models.
- **2026** — Co-authored the persona selection model essay, contributed to Claude's Constitution, and spoke at the Vatican presentation of Pope Leo XIV's AI encyclical.

## Achievements

- Helped establish mechanistic interpretability as a distinct approach to reverse-engineering learned neural-network computations.
- Co-founded Anthropic in 2021 and leads its interpretability research.
- Co-founded Distill, which made interactive visualization and clear exposition central to communicating machine-learning research.
- Contributed to DeepDream, feature visualization, activation atlases, the OpenAI Circuits program, and Anthropic's feature and circuit tracing work.
- Co-authored Concrete Problems in AI Safety and multiple widely cited interpretability papers and essays.
- Selected for the TIME100 AI list in 2024.
- Received a Thiel Fellowship in 2012.

## Latest updates

- **2026-05** — Spoke at the Vatican presentation of Pope Leo XIV's AI encyclical, calling for independent moral voices and broader social participation in shaping AI.
- **2026-05** — Anthropic published Natural Language Autoencoders, part of the interpretability team's effort to translate model activations into human-readable descriptions.
- **2026-02** — Co-authored The Persona Selection Model, an essay arguing that AI assistants can be understood partly as learned character-like personas elicited during post-training.
- **2026-01** — Contributed substantial material on model nature, identity, and psychology to Claude's Constitution.
- **2025-10** — Co-authored research on how language models manipulate geometric representations while solving a text-counting task.
- **2025-08** — Published an informal note on mechanistic faithfulness and the risk that sparse approximations may implement different mechanisms from the models they approximate.

## Links

- Website: https://colah.github.io/
- Twitter/X: https://x.com/ch402
- GitHub: https://github.com/colah
- Bluesky: https://bsky.app/profile/colah.bsky.social

---

Profile page: https://yespress.io/chris-olah
Published by YesPress — https://yespress.io
Last updated: 2026-09-26
