distillation
15 posts tagged distillation.
- UECR-GRPO treats the teacher as evidence, not an oracle
- Xiaomi's MiMo V2.6 Ships in Three Flavors, and the Split Matters
- When the Teacher Knows to Quit: RetireOPD and Distillation for Agents
- Distillation needs calibration when the teacher is biased
- OptiFlow treats offline RL policy learning as sample matching
- Garry Tan’s distillation argument is really about AI capability access
- On-policy distillation may need fewer prompts and better absorption
- On-policy distillation may be pruning tails, not teaching
- The token-budget bug hiding in multi-teacher distillation
- When the Teacher and the Verifier Disagree: Fixing On-Policy Distillation for Long Context
- A Strong Model Can Scaffold a Weak One Without Any Retraining
- Relay-OPD and the prefix failure problem in on-policy distillation
- OPD2 tries to distill reasoning by subtracting the base model
- Reusing RL Gains Across Model Sizes: The Case for Direct-OPD
- DemoPSD Treats Teacher Disagreement as a Training Signal