Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

post-training

20 posts tagged post-training.

  • Teaching a Reasoning Model to Know When It's Sure Cuts Its Token Bill Sep 28, 2026
  • MISVO steers frozen language models at inference time Sep 25, 2026
  • The World Model That Edits the Agent's Mind, Not the Environment Sep 24, 2026
  • onPanda turns alignment feedback into token-level steering Sep 22, 2026
  • Jev is a decision model, not another chatbot Sep 18, 2026
  • When the Teacher Knows to Quit: RetireOPD and Distillation for Agents Sep 18, 2026
  • ComPO and the Case Against Optimizing the Loss You Wrote Down Sep 17, 2026
  • RetroThinker lets speech models correct themselves while you are still talking Sep 11, 2026
  • Layer-selective unlearning aims at the parts of a model that remember Sep 10, 2026
  • An AI Just Outscored the Top Human at IOI 2026. Here's What That Actually Means Sep 3, 2026
  • The SFT-RL Split Has a Wide Safe Zone, and You Can Find It Cheap Sep 2, 2026
  • Small dialogue agents need repair loops, not just bigger training Aug 31, 2026
  • IAR makes retrieval-free document QA less brittle Aug 21, 2026
  • The token-budget bug hiding in multi-teacher distillation Aug 20, 2026
  • When the Teacher and the Verifier Disagree: Fixing On-Policy Distillation for Long Context Aug 20, 2026
  • Post-training is now the behavior layer Aug 3, 2026
  • Output Reset makes PPO smoother, not automatically better Jul 21, 2026
  • Muon’s agent RL win is real, narrow, and stack-dependent Jul 20, 2026
  • Reusing RL Gains Across Model Sizes: The Case for Direct-OPD Jul 7, 2026
  • Post-training logprobs as a step-level critic for agents Jun 25, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public