Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

reasoning-models

31 posts tagged reasoning-models.

  • Teaching a Reasoning Model to Know When It's Sure Cuts Its Token Bill Sep 28, 2026
  • ExplorationBench and the Gap Between Recall and Real Discovery Sep 25, 2026
  • The Modality Gap in Speech Fact-Checking, and Why Retrieval Alone Doesn't Fix It Sep 25, 2026
  • Clifford-VAE puts pixels into symbolic memory space Sep 24, 2026
  • SWE-Flux: The Benchmark That Asks If Coding Models Can Predict What Code Actually Does Sep 24, 2026
  • The World Model That Edits the Agent's Mind, Not the Environment Sep 24, 2026
  • UECR-GRPO treats the teacher as evidence, not an oracle Sep 24, 2026
  • When a Model Gets Math Right but Reads the Same Problem Differently Sep 24, 2026
  • A Small Model That Steers a Bigger One's Reasoning Sep 23, 2026
  • OOD generalization depends on exact mechanisms, not better fits Sep 22, 2026
  • OpenAI’s math advisory group is about claim control Sep 21, 2026
  • A cipher win is not an eval without the working Sep 19, 2026
  • Bellman Policy Optimization cuts one moving part from RLVR Sep 15, 2026
  • AI math systems are still optimizing for the wrong proof Sep 12, 2026
  • RetroThinker lets speech models correct themselves while you are still talking Sep 11, 2026
  • ConvMem Turns Long-Context Reasoning Into a Tree Instead of a Chain Sep 10, 2026
  • Geometry reasoning gets better when the model is not doing every job Sep 10, 2026
  • OpenAI’s reported Millennium math claim needs proof, not applause Sep 9, 2026
  • An AI Just Outscored the Top Human at IOI 2026. Here's What That Actually Means Sep 3, 2026
  • Prefix Sliding bets most reasoning tokens are dead weight Aug 27, 2026
  • BDH-CQ Makes ARC Reasoning Cheaper by Thinking in Latent Space Aug 11, 2026
  • Concise answers can weaken reasoning in fused LLM training Aug 11, 2026
  • What 'Test-Time Scaling' Actually Means When You Read a Benchmark Aug 5, 2026
  • AI reasoning can look right while taking shortcuts Aug 1, 2026
  • OpenAI's Ten Math Results: What Counts as a Real Advance Aug 1, 2026
  • PPL-Factory makes the case for smaller fine-tuning sets Jul 21, 2026
  • Chess Shows What RL Actually Does to a Reasoning Model Jul 20, 2026
  • OPD2 tries to distill reasoning by subtracting the base model Jul 17, 2026
  • AdaPrefix-GRPO turns hard reasoning failures into training signal Jul 9, 2026
  • Agon Grades the Reasoning, Not Just the Answer Jul 9, 2026
  • Reasoning Traces as a Difficulty Sensor: What Epi2Diff Gets Right Jun 29, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public