reinforcement-learning
24 posts tagged reinforcement-learning.
- RCI turns stop signals into safer offline RL training data
- When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL
- The Alignment Tax on Creativity, and a Switch to Turn It Back On
- Moment closure brings uncertainty back into model-based RL planning
- Two prepared policies may be the sweet spot for uncertain MDPs
- RL can teach code models to care about runtime, but the stopwatch is the hard part
- Training Agents Inside the Harness They Actually Ship With
- Chess Shows What RL Actually Does to a Reasoning Model
- DADiff uses diffusion to measure when an RL policy stops transferring
- Teaching a Model to Zoom In on a Chart Before It Judges a Claim
- The Missing Half of RL for Diffusion Language Models
- TerraZero bets on self-play for the driving long tail
- REGRIND’s one-demo recipe for robot hands
- Fraud detection needs a Pareto frontier, not another magic score
- Agon Grades the Reasoning, Not Just the Answer
- Reusing RL Gains Across Model Sizes: The Case for Direct-OPD
- Verification as a Scaling Axis: What the LLM-as-a-Verifier Paper Actually Changes
- LeRobot v0.6.0 Turns Robot Simulation Into a Feedback Loop
- The Cheap Way to Test Whether Your Agent's Step Scores Actually Mean Anything
- TRIAGE gives agent RL a better target than pass or fail
- Finger-Level Ownership: How DexCompose Stacks Robot Hand Skills Without Breaking Them
- HiReLC Treats Model Compression as Coordination, Not a Knob
- Tool-use RL is failing at the brackets, not the tools
- CoorDex and the End of Stop-and-Go Humanoids