Media Summary: [PoD] Self-Distilled Reasoner : On-Policy Self-Distillation for Large Language Models I recently met Sasha Rush and he started giving me an impromptu lecture on how targeted on- How do we train a small yet powerful AI model? This video explains knowledge

Self Distilled Reasoner On Policy - Detailed Analysis & Overview

[PoD] Self-Distilled Reasoner : On-Policy Self-Distillation for Large Language Models I recently met Sasha Rush and he started giving me an impromptu lecture on how targeted on- How do we train a small yet powerful AI model? This video explains knowledge Disclaimer: This video is generated with Google's NotebookLM. Rethinking On- In this AI Research Roundup episode, Alex discusses the paper: ' In this video, we sit down with Jonas Hübotter (ETH Zurich) and Idan Shenfeld (MIT) to break down

00:00 - Intro 01:39 - Motivation: Why On-

Photo Gallery

What is the Self-Distilled Reasoner? On-Policy Self-Distillation
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models (Jan 2026)
On-Policy Self-Distillation for Reasoning Compression (Mar 2026)
[PoD] Self-Distilled Reasoner : On-Policy Self-Distillation for Large Language Models
How On Policy Self Distillation Works - Sasha Rush
What If Distillation Could Learn Like RL?
Fast and Effective On-policy Distillation from Reasoning Prefixes
Rethinking On-Policy Distillation of Large Language Models
Self-Distilled RLVR: Stable LLM Training Method
Why Self-Distillation Is Taking Over LLM Post-Training (w/ the Researchers Behind It)
On-Policy Distillation in 20 Min
Self-Distilled Reasoner: Teaching Models to Learn from Themselves
View Detailed Profile
What is the Self-Distilled Reasoner? On-Policy Self-Distillation

What is the Self-Distilled Reasoner? On-Policy Self-Distillation

What is the

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models (Jan 2026)

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models (Jan 2026)

Title:

On-Policy Self-Distillation for Reasoning Compression (Mar 2026)

On-Policy Self-Distillation for Reasoning Compression (Mar 2026)

Title: On-

[PoD] Self-Distilled Reasoner : On-Policy Self-Distillation for Large Language Models

[PoD] Self-Distilled Reasoner : On-Policy Self-Distillation for Large Language Models

[PoD] Self-Distilled Reasoner : On-Policy Self-Distillation for Large Language Models

How On Policy Self Distillation Works - Sasha Rush

How On Policy Self Distillation Works - Sasha Rush

I recently met Sasha Rush and he started giving me an impromptu lecture on how targeted on-

What If Distillation Could Learn Like RL?

What If Distillation Could Learn Like RL?

How do we train a small yet powerful AI model? This video explains knowledge

Fast and Effective On-policy Distillation from Reasoning Prefixes

Fast and Effective On-policy Distillation from Reasoning Prefixes

Paper: Fast and Effective On-

Rethinking On-Policy Distillation of Large Language Models

Rethinking On-Policy Distillation of Large Language Models

Disclaimer: This video is generated with Google's NotebookLM. Rethinking On-

Self-Distilled RLVR: Stable LLM Training Method

Self-Distilled RLVR: Stable LLM Training Method

In this AI Research Roundup episode, Alex discusses the paper: '

Why Self-Distillation Is Taking Over LLM Post-Training (w/ the Researchers Behind It)

Why Self-Distillation Is Taking Over LLM Post-Training (w/ the Researchers Behind It)

In this video, we sit down with Jonas Hübotter (ETH Zurich) and Idan Shenfeld (MIT) to break down

On-Policy Distillation in 20 Min

On-Policy Distillation in 20 Min

00:00 - Intro 01:39 - Motivation: Why On-

Self-Distilled Reasoner: Teaching Models to Learn from Themselves

Self-Distilled Reasoner: Teaching Models to Learn from Themselves

Paper:

Self-Distilled Policy Gradient

Self-Distilled Policy Gradient

ai #research https://arxiv.org/pdf/2606.04036