Media Summary: Authors: Chen, Wei-Chi; Chu, Wei-Ta* Description: With labeled data, Continual learning is an important component of AGI and ASI, and this paper presents a new idea in this area using ... Papers: DINO: DINOv2: ViTs Need Registers: ...

W07 2 Self Distillation Visual - Detailed Analysis & Overview

Authors: Chen, Wei-Chi; Chu, Wei-Ta* Description: With labeled data, Continual learning is an important component of AGI and ASI, and this paper presents a new idea in this area using ... Papers: DINO: DINOv2: ViTs Need Registers: ... Fine-tuning a diffusion model usually means running reinforcement learning for days, which costs a lot of GPU-hours. ByteDance's ... This week we review the paper Reinforcement Learning via This research introduces a groundbreaking framework for test-time

Photo Gallery

W07.2: Self-distillation visual representation learning without negatives: BYOL, DINO (Part 2/2)
W07.1: Self-distillation visual representation learning without negatives: BYOL, DINO (Part 1/2)
SSSD: Self-Supervised Self Distillation
Self-Distillation Enables Continual Learning
Self-Distillation Enables Continual Learning Paper-2026
Paper: Visual Contrastive Self-Distillation
Self Supervised Learning through self distillation with no labels, and Why ViTs Need Registers
No RL, 40–63% Cheaper: How Diffusion Models Learn From Themselves
2601.19897 - Self-Distillation Enables Continual Learning
Self-Distillation as a New Framework for Continual Learning | Idan Shenfeld | Random Samples
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models (Jan 2026)
Reinforcement Learning via Self-Distillation
View Detailed Profile
W07.2: Self-distillation visual representation learning without negatives: BYOL, DINO (Part 2/2)

W07.2: Self-distillation visual representation learning without negatives: BYOL, DINO (Part 2/2)

Part

W07.1: Self-distillation visual representation learning without negatives: BYOL, DINO (Part 1/2)

W07.1: Self-distillation visual representation learning without negatives: BYOL, DINO (Part 1/2)

Part 1/

SSSD: Self-Supervised Self Distillation

SSSD: Self-Supervised Self Distillation

Authors: Chen, Wei-Chi; Chu, Wei-Ta* Description: With labeled data,

Self-Distillation Enables Continual Learning

Self-Distillation Enables Continual Learning

Paper:

Self-Distillation Enables Continual Learning Paper-2026

Self-Distillation Enables Continual Learning Paper-2026

Continual learning is an important component of AGI and ASI, and this paper presents a new idea in this area using ...

Paper: Visual Contrastive Self-Distillation

Paper: Visual Contrastive Self-Distillation

On-policy

Self Supervised Learning through self distillation with no labels, and Why ViTs Need Registers

Self Supervised Learning through self distillation with no labels, and Why ViTs Need Registers

Papers: DINO: https://arxiv.org/abs/2104.14294 DINOv2: https://arxiv.org/abs/2304.07193 ViTs Need Registers: ...

No RL, 40–63% Cheaper: How Diffusion Models Learn From Themselves

No RL, 40–63% Cheaper: How Diffusion Models Learn From Themselves

Fine-tuning a diffusion model usually means running reinforcement learning for days, which costs a lot of GPU-hours. ByteDance's ...

2601.19897 - Self-Distillation Enables Continual Learning

2601.19897 - Self-Distillation Enables Continual Learning

title:

Self-Distillation as a New Framework for Continual Learning | Idan Shenfeld | Random Samples

Self-Distillation as a New Framework for Continual Learning | Idan Shenfeld | Random Samples

Self

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models (Jan 2026)

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models (Jan 2026)

Title:

Reinforcement Learning via Self-Distillation

Reinforcement Learning via Self-Distillation

This week we review the paper Reinforcement Learning via

Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation

Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation

This research introduces a groundbreaking framework for test-time