Media Summary: MIT 6.S897 Machine Learning for Healthcare, Spring 2019 Instructor: Peter Szolovits View the complete course: ... Lex Fridman Podcast full episode: Thank you for listening ❤ Check out our ... Adam Shai presented “Building the Science of

25 Interpretability - Detailed Analysis & Overview

MIT 6.S897 Machine Learning for Healthcare, Spring 2019 Instructor: Peter Szolovits View the complete course: ... Lex Fridman Podcast full episode: Thank you for listening ❤ Check out our ... Adam Shai presented “Building the Science of What's happening inside an AI model as it thinks? Why are AI models sycophantic, and why do they hallucinate? Are AI models ... How can we reverse engineer what a neural network is doing? In this IASEAI ' A surprising fact about modern large language models is that nobody really knows how they work internally. At Anthropic, the ...

Recording of the webinar from 2021-11-19 0:00 Introductory Remarks 0:50 Seminar Overview 1:41 An Abundance of Data 2:53 ... Chenhao Tan demonstrates an automated mechanistic Christoph Molnar is one of the main people to know in the space of When Anthropic tested Claude Sonnet 4.5 for alignment, the model appeared perfectly behaved — but it turned out the model had ...

Photo Gallery

25. Interpretability
Mechanistic Interpretability explained | Chris Olah and Lex Fridman
Adam Shai - Building the Science of Interpretability
Interpretability: Understanding how AI models think
Guide Labs: Why AI Interpretability Has to Start at Training Time
An Introduction to Mechanistic Interpretability – Neel Nanda | IASEAI 2025
What is interpretability?
Interpretable ML: SISSO, XGBoost, TPOT, Roost, Webinar 2021-11-19
Chenhao Tan - Automating Mechanistic Interpretability [Alignment Workshop]
#047 Interpretable Machine Learning - Christoph Molnar
Interpretable vs Explainable Machine Learning
Neel Nanda - Our Pivot To Pragmatic Interpretability [Alignment Workshop]
View Detailed Profile
25. Interpretability

25. Interpretability

MIT 6.S897 Machine Learning for Healthcare, Spring 2019 Instructor: Peter Szolovits View the complete course: ...

Mechanistic Interpretability explained | Chris Olah and Lex Fridman

Mechanistic Interpretability explained | Chris Olah and Lex Fridman

Lex Fridman Podcast full episode: https://www.youtube.com/watch?v=ugvHCXCOmm4 Thank you for listening ❤ Check out our ...

Adam Shai - Building the Science of Interpretability

Adam Shai - Building the Science of Interpretability

Adam Shai presented “Building the Science of

Interpretability: Understanding how AI models think

Interpretability: Understanding how AI models think

What's happening inside an AI model as it thinks? Why are AI models sycophantic, and why do they hallucinate? Are AI models ...

Guide Labs: Why AI Interpretability Has to Start at Training Time

Guide Labs: Why AI Interpretability Has to Start at Training Time

Much of AI

An Introduction to Mechanistic Interpretability – Neel Nanda | IASEAI 2025

An Introduction to Mechanistic Interpretability – Neel Nanda | IASEAI 2025

How can we reverse engineer what a neural network is doing? In this IASEAI '

What is interpretability?

What is interpretability?

A surprising fact about modern large language models is that nobody really knows how they work internally. At Anthropic, the ...

Interpretable ML: SISSO, XGBoost, TPOT, Roost, Webinar 2021-11-19

Interpretable ML: SISSO, XGBoost, TPOT, Roost, Webinar 2021-11-19

Recording of the webinar from 2021-11-19 0:00 Introductory Remarks 0:50 Seminar Overview 1:41 An Abundance of Data 2:53 ...

Chenhao Tan - Automating Mechanistic Interpretability [Alignment Workshop]

Chenhao Tan - Automating Mechanistic Interpretability [Alignment Workshop]

Chenhao Tan demonstrates an automated mechanistic

#047 Interpretable Machine Learning - Christoph Molnar

#047 Interpretable Machine Learning - Christoph Molnar

Christoph Molnar is one of the main people to know in the space of

Interpretable vs Explainable Machine Learning

Interpretable vs Explainable Machine Learning

Interpretable

Neel Nanda - Our Pivot To Pragmatic Interpretability [Alignment Workshop]

Neel Nanda - Our Pivot To Pragmatic Interpretability [Alignment Workshop]

When Anthropic tested Claude Sonnet 4.5 for alignment, the model appeared perfectly behaved — but it turned out the model had ...

Pop Goes the Stack | Mechanistic Interpretability: Debugging LLMs by reading their circuits | AI

Pop Goes the Stack | Mechanistic Interpretability: Debugging LLMs by reading their circuits | AI

Mechanistic