reasoning
9 posts tagged reasoning.
- DungeonBench puts tactical reasoning where agents usually break
- Relay-OPD and the prefix failure problem in on-policy distillation
- MIRROR trains vision models by making each modality teach the others
- PoTRE makes the case for heterogeneous test-time reasoning
- Goal prompting is not a solver for NP-hard search
- The Missing Half of RL for Diffusion Language Models
- RL post-training as procedure compression, not just skill amplification
- DemoPSD Treats Teacher Disagreement as a Training Signal
- Theoria Makes a Model Show Its Work, Then Checks Every Line