reinforcement-learning
45 posts tagged reinforcement-learning.
- UECR-GRPO treats the teacher as evidence, not an oracle
- A Small Model That Steers a Bigger One's Reasoning
- Speaker-Centered Memory: Why Group Chats Break Your AI Agent
- Critical-State RL trains the one agent call that actually matters
- Where Harness Self-Improvement Actually Stands in Late 2026
- Xiaomi's MiMo V2.6 Ships in Three Flavors, and the Split Matters
- When the Teacher Knows to Quit: RetireOPD and Distillation for Agents
- ScienceBuddy and the case for training the harness before the model
- Bellman Policy Optimization cuts one moving part from RLVR
- OptiFlow treats offline RL policy learning as sample matching
- Recursive Self-Improvement Has a Roadmap Now, and Most of It Isn't Built
- SAGE Uses a Big Model as a Coach, Not a Crutch
- The SFT-RL Split Has a Wide Safe Zone, and You Can Find It Cheap
- Verbal reinforcement learning is a feedback routing problem
- Aero Hand Open makes the cheap part of a robot hand the learnable part
- Can a robot think out loud before it moves? R³ tests the idea
- BPCO Brings the Critic Back to RL Fine-Tuning
- Splitting the reward: how G-CARL grades medical explanations for facts and for tone
- PGFS++ and the Reward Magnet Problem in AI Drug Design
- SPADE Makes the Training Environment a Thing the Model Learns to Build
- LLM reward shaping without changing the goal
- RCI turns stop signals into safer offline RL training data
- When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL
- The Alignment Tax on Creativity, and a Switch to Turn It Back On
- Moment closure brings uncertainty back into model-based RL planning
- Two prepared policies may be the sweet spot for uncertain MDPs
- RL can teach code models to care about runtime, but the stopwatch is the hard part
- Training Agents Inside the Harness They Actually Ship With
- Chess Shows What RL Actually Does to a Reasoning Model
- DADiff uses diffusion to measure when an RL policy stops transferring
- Teaching a Model to Zoom In on a Chart Before It Judges a Claim
- The Missing Half of RL for Diffusion Language Models
- TerraZero bets on self-play for the driving long tail
- REGRIND’s one-demo recipe for robot hands
- Fraud detection needs a Pareto frontier, not another magic score
- Agon Grades the Reasoning, Not Just the Answer
- Reusing RL Gains Across Model Sizes: The Case for Direct-OPD
- Verification as a Scaling Axis: What the LLM-as-a-Verifier Paper Actually Changes
- LeRobot v0.6.0 Turns Robot Simulation Into a Feedback Loop
- The Cheap Way to Test Whether Your Agent's Step Scores Actually Mean Anything
- TRIAGE gives agent RL a better target than pass or fail
- Finger-Level Ownership: How DexCompose Stacks Robot Hand Skills Without Breaking Them
- HiReLC Treats Model Compression as Coordination, Not a Knob
- Tool-use RL is failing at the brackets, not the tools
- CoorDex and the End of Stop-and-Go Humanoids