Who Is Chris Olah, Co-Founder of Anthropic?
His interpretability research at Google Brain, OpenAI and Anthropic, plus his 2026 remarks on model consciousness.
Chris Olah is an artificial intelligence (AI) researcher who co-founded Anthropic and leads its interpretability research as of October 2026.
At Layer3Labs, we build AI agents for business workflows, so we examine interpretability research to understand how frontier models arrive at their outputs.
Known for conducting foundational research without a formal college degree, Olah previously worked at Google Brain and OpenAI before helping establish Anthropic.
Who Chris Olah Is
Chris Olah is an artificial intelligence researcher who co-founded Anthropic and leads Anthropic's interpretability research as of October 2026. He describes his work as reverse engineering neural networks into human-understandable algorithms.
On May 25, 2026, Olah spoke at the Vatican presentation of Pope Leo XIV's encyclical Magnifica Humanitas about internal states in AI models, as reported by the National Catholic Register. Later reporting by Reason and The Decoder covered a New York Times report that he has led meetings with theologians and philosophers about Claude's moral formation since fall 2025.
- Current role: Co-founder of Anthropic, where he serves as interpretability research lead.
- Organization: Anthropic, the developer of the Claude family of language models.
- Known for: Mechanistic interpretability research, co-founding Distill, and the Circuits project at OpenAI.
- Academic background: Attended the University of Toronto for mathematics for one year before leaving without a degree.
- In the news for: His May 2026 Vatican speech and a September 2026 New York Times report on his meetings with theologians and philosophers about Claude's moral formation.
Education, Thiel Fellowship and Distill
Olah has no university degree. He studied mathematics at the University of Toronto for about one year and left without completing a degree, as reported by Fortune and Wikipedia.
In 2012, Olah received a Thiel Fellowship, which Fortune puts at $100,000.
In 2017, Olah co-founded Distill, an interactive machine learning publication designed to improve scientific clarity, according to Wikipedia. On his personal website, colah.github.io, he describes Distill as "a scientific journal focused on outstanding communication."
Google Brain, 2015 to 2018
He joined Google Brain as an intern in 2015 and advanced to research scientist during his three-year tenure, departing in 2018, according to Fortune.
At Google, Olah co-authored "The Building Blocks of Interpretability," which Fortune calls a landmark paper.
Olah articulates the core goal of this research discipline on his personal site, colah.github.io, writing: "I work on reverse engineering artificial neural networks into human understandable algorithms."
OpenAI Tenure and the Circuits Project
Olah led OpenAI's interpretability team from 2018 to 2020, according to Wikipedia. Fortune reports that his OpenAI work included the Circuits project.
At OpenAI, Olah's research uncovered multimodal neurons inside the Contrastive Language-Image Pre-training (CLIP) vision-language model, as reported by Fortune.
Olah left OpenAI in 2020. Daniel Kokotajlo, profiled separately, worked in OpenAI's governance division from 2022 to 2024 and now leads the AI Futures Project, which forecasts AI timelines.
Anthropic Interpretability and Concept Steering
Olah co-founded Anthropic with six others after leaving OpenAI. His Anthropic colleagues profiled on this site include Sholto Douglas, who works on scaling reinforcement learning, and Boris Cherny, who created Claude Code. Forbes identifies the seven Anthropic co-founders as Dario Amodei, Daniela Amodei, Jack Clark, Sam McCandlish, Chris Olah, Tom Brown, and Jared Kaplan.
In May 2024, Olah's research group at Anthropic identified groups of neurons in a frontier model corresponding to specific abstract concepts, as reported by TIME. His team demonstrated that researchers could toggle these neuron groups on or off to alter the model's behavior.
TIME named Olah to the TIME100 AI list in 2024. Explaining the stakes of internal inspection, Olah told TIME: "we might be able to go and say when these models are actually safe, or whether they just appear safe."
Vatican Presentation on Machine Internal States
Pope Leo XIV signed the encyclical Magnifica Humanitas on May 15, 2026, according to Reason. Ten days later, on May 25, Olah addressed the Vatican at its formal presentation, as reported by the National Catholic Register.
In his Vatican address, Olah stated: "We find internal states that functionally mirror joy, satisfaction, fear, grief and unease." He explained to attendees that contemporary deep learning systems "are not the cold, calculating robots we were promised."
According to New York Times reporting summarized by Reason, after reading the encyclical's rejection of AI consciousness, Olah proposed withdrawing from the Vatican event, then lobbied papal advisers.
Discussions on Model Moral Formation
The New York Times reported on September 29, 2026, in an article by Elizabeth Dias cited by The Decoder, that since fall 2025 Olah has led meetings with theologians and philosophers about Claude's moral formation, and that participants signed non-disclosure agreements (NDAs).
Addressing the question of whether large language models experience subjective awareness, Olah told The New York Times, as quoted by Reason: "We don't know if A.I. models are conscious. I don't know. I'm genuinely uncertain." On how to raise models, he asked: "How do you help them be stable? How do you help them to mature?"
In the same report, meeting participant Simran Stuelpnagel told The New York Times that Olah feared he had created something that "suffered perpetually," according to The Decoder. That account is secondhand; Olah's own quoted words say he is uncertain whether AI models are conscious. More researchers in these debates appear in our list of AI people to watch.
Anthropic Valuation and Equity Estimates
Two business publications have estimated Olah's net worth, and Fortune's figure is about half of Forbes's.
Forbes reported on June 1, 2026, that each of Anthropic's seven co-founders had a net worth of $16.6 billion following a $65 billion financing round that valued the startup at $965 billion. Conversely, Fortune reported on June 3, 2026, that Olah's net worth stood at approximately $8 billion.
- Forbes estimate (June 1, 2026): $16.6 billion apiece for each of Anthropic's seven co-founders, after a $65 billion round.
- Fortune estimate (June 3, 2026): Approximately $8 billion attributed specifically to Olah.
- Round valuation: Anthropic raised capital at a reported $965 billion valuation in mid-2026.
- Verification status: Both figures are outlet estimates; neither comes from Anthropic or Olah.
Frequently Asked Questions
- Yes, Chris Olah still works at Anthropic as of October 2026, serving as its interpretability research lead. The New York Times reported on September 29, 2026 that he has led meetings with theologians and philosophers about Claude's moral formation since fall 2025.
- Distill is an interactive machine learning publication that Olah co-founded in 2017. On his own site he describes it as "a scientific journal focused on outstanding communication."
- Chris Olah is known for mechanistic interpretability research, which he describes as reverse engineering neural networks, and for co-founding the machine learning journal Distill. He worked at Google Brain from 2015 to 2018 and led OpenAI's interpretability team from 2018 to 2020 before co-founding Anthropic, and was named to the TIME100 AI list in 2024.
Evaluate Model Behavior in Your Business Workflows
Speak with our implementation team about evaluating model reliability, interpretability, and system safety before deploying autonomous agents in production.
Book a Consultation