model-training
13 posts tagged model-training.
- Alignment baked into pretraining, not bolted on later
- Selective Trust: Why RAG Systems Should Learn When to Ignore Their Own Context
- MindForge trains coding agents on blank-repo software work
- Hyperball optimizers still need learning-rate discipline
- MIRROR trains vision models by making each modality teach the others
- OPD2 tries to distill reasoning by subtracting the base model
- RL post-training as procedure compression, not just skill amplification
- Timestamp drift is the quiet ASR failure that breaks real workflows
- DemoPSD Treats Teacher Disagreement as a Training Signal
- RLVR needs a taste model, not just a grader
- Synthetic QA Has a Selection Problem Before It Has a Training Problem
- Autodata Turns Synthetic Data Generation Into an Agent You Train
- Self-distillation can make models better on the first try and worse on the fifth