Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

reinforcement-learning

45 posts tagged reinforcement-learning.

  • UECR-GRPO treats the teacher as evidence, not an oracle Sep 24, 2026
  • A Small Model That Steers a Bigger One's Reasoning Sep 23, 2026
  • Speaker-Centered Memory: Why Group Chats Break Your AI Agent Sep 23, 2026
  • Critical-State RL trains the one agent call that actually matters Sep 22, 2026
  • Where Harness Self-Improvement Actually Stands in Late 2026 Sep 22, 2026
  • Xiaomi's MiMo V2.6 Ships in Three Flavors, and the Split Matters Sep 22, 2026
  • When the Teacher Knows to Quit: RetireOPD and Distillation for Agents Sep 18, 2026
  • ScienceBuddy and the case for training the harness before the model Sep 16, 2026
  • Bellman Policy Optimization cuts one moving part from RLVR Sep 15, 2026
  • OptiFlow treats offline RL policy learning as sample matching Sep 15, 2026
  • Recursive Self-Improvement Has a Roadmap Now, and Most of It Isn't Built Sep 11, 2026
  • SAGE Uses a Big Model as a Coach, Not a Crutch Sep 2, 2026
  • The SFT-RL Split Has a Wide Safe Zone, and You Can Find It Cheap Sep 2, 2026
  • Verbal reinforcement learning is a feedback routing problem Sep 2, 2026
  • Aero Hand Open makes the cheap part of a robot hand the learnable part Aug 31, 2026
  • Can a robot think out loud before it moves? R³ tests the idea Aug 27, 2026
  • BPCO Brings the Critic Back to RL Fine-Tuning Aug 25, 2026
  • Splitting the reward: how G-CARL grades medical explanations for facts and for tone Aug 21, 2026
  • PGFS++ and the Reward Magnet Problem in AI Drug Design Aug 20, 2026
  • SPADE Makes the Training Environment a Thing the Model Learns to Build Aug 20, 2026
  • LLM reward shaping without changing the goal Aug 19, 2026
  • RCI turns stop signals into safer offline RL training data Aug 13, 2026
  • When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL Aug 13, 2026
  • The Alignment Tax on Creativity, and a Switch to Turn It Back On Aug 10, 2026
  • Moment closure brings uncertainty back into model-based RL planning Aug 4, 2026
  • Two prepared policies may be the sweet spot for uncertain MDPs Aug 4, 2026
  • RL can teach code models to care about runtime, but the stopwatch is the hard part Jul 29, 2026
  • Training Agents Inside the Harness They Actually Ship With Jul 24, 2026
  • Chess Shows What RL Actually Does to a Reasoning Model Jul 20, 2026
  • DADiff uses diffusion to measure when an RL policy stops transferring Jul 20, 2026
  • Teaching a Model to Zoom In on a Chart Before It Judges a Claim Jul 20, 2026
  • The Missing Half of RL for Diffusion Language Models Jul 17, 2026
  • TerraZero bets on self-play for the driving long tail Jul 15, 2026
  • REGRIND’s one-demo recipe for robot hands Jul 14, 2026
  • Fraud detection needs a Pareto frontier, not another magic score Jul 13, 2026
  • Agon Grades the Reasoning, Not Just the Answer Jul 9, 2026
  • Reusing RL Gains Across Model Sizes: The Case for Direct-OPD Jul 7, 2026
  • Verification as a Scaling Axis: What the LLM-as-a-Verifier Paper Actually Changes Jul 7, 2026
  • LeRobot v0.6.0 Turns Robot Simulation Into a Feedback Loop Jul 6, 2026
  • The Cheap Way to Test Whether Your Agent's Step Scores Actually Mean Anything Jul 1, 2026
  • TRIAGE gives agent RL a better target than pass or fail Jul 1, 2026
  • Finger-Level Ownership: How DexCompose Stacks Robot Hand Skills Without Breaking Them Jun 29, 2026
  • HiReLC Treats Model Compression as Coordination, Not a Knob Jun 25, 2026
  • Tool-use RL is failing at the brackets, not the tools Jun 25, 2026
  • CoorDex and the End of Stop-and-Go Humanoids Jun 23, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public