post-training
20 posts tagged post-training.
- Teaching a Reasoning Model to Know When It's Sure Cuts Its Token Bill
- MISVO steers frozen language models at inference time
- The World Model That Edits the Agent's Mind, Not the Environment
- onPanda turns alignment feedback into token-level steering
- Jev is a decision model, not another chatbot
- When the Teacher Knows to Quit: RetireOPD and Distillation for Agents
- ComPO and the Case Against Optimizing the Loss You Wrote Down
- RetroThinker lets speech models correct themselves while you are still talking
- Layer-selective unlearning aims at the parts of a model that remember
- An AI Just Outscored the Top Human at IOI 2026. Here's What That Actually Means
- The SFT-RL Split Has a Wide Safe Zone, and You Can Find It Cheap
- Small dialogue agents need repair loops, not just bigger training
- IAR makes retrieval-free document QA less brittle
- The token-budget bug hiding in multi-teacher distillation
- When the Teacher and the Verifier Disagree: Fixing On-Policy Distillation for Long Context
- Post-training is now the behavior layer
- Output Reset makes PPO smoother, not automatically better
- Muon’s agent RL win is real, narrow, and stack-dependent
- Reusing RL Gains Across Model Sizes: The Case for Direct-OPD
- Post-training logprobs as a step-level critic for agents