Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

ai-research

75 posts tagged ai-research.

  • Flow matching gives neural dissimilarity metrics one common frame Sep 28, 2026
  • Teaching a Reasoning Model to Know When It's Sure Cuts Its Token Bill Sep 28, 2026
  • What LLMs Miss About Haitian Creole, and Why Low-Resource Culture Breaks Evals Sep 28, 2026
  • Conspiracy Detection Needs Context, Not Just Better Keywords Sep 25, 2026
  • ExplorationBench and the Gap Between Recall and Real Discovery Sep 25, 2026
  • MISVO steers frozen language models at inference time Sep 25, 2026
  • Authorship verification works better as similarity, not classification Sep 24, 2026
  • Clifford-VAE puts pixels into symbolic memory space Sep 24, 2026
  • When a Model Gets Math Right but Reads the Same Problem Differently Sep 24, 2026
  • A Small Model That Steers a Bigger One's Reasoning Sep 23, 2026
  • Long-context models can miss the answer sitting behind nearby noise Sep 23, 2026
  • OOD generalization depends on exact mechanisms, not better fits Sep 22, 2026
  • Membership Inference Gets an Entropy Correction: Reading the ETD Paper Sep 21, 2026
  • OpenAI’s math advisory group is about claim control Sep 21, 2026
  • Softmax attention needs an off switch Sep 21, 2026
  • Decontamination reports are a promise, not a proof: the case for evaluation-side benchmarks Sep 20, 2026
  • A cipher win is not an eval without the working Sep 19, 2026
  • ComPO and the Case Against Optimizing the Loss You Wrote Down Sep 17, 2026
  • Double descent as an implicit regularization story Sep 17, 2026
  • Probabilistic Linear Explanations make interpretability more honest Sep 17, 2026
  • What Actually Makes a Tokeniser Good: Search Beats Objective Sep 17, 2026
  • When Rewording the Answer Key Reshuffles the Leaderboard Sep 17, 2026
  • Fuse tests the weak spot in AI social advice Sep 16, 2026
  • LACE compresses speech tokens one codec layer at a time Sep 16, 2026
  • ScienceBuddy and the case for training the harness before the model Sep 16, 2026
  • Bellman Policy Optimization cuts one moving part from RLVR Sep 15, 2026
  • LLM personas fail when opinions have to change Sep 15, 2026
  • OptiFlow treats offline RL policy learning as sample matching Sep 15, 2026
  • Slip Detection Is Where Robot Hands Stop Dropping Things Sep 15, 2026
  • The Skill Router You Already Have: Gavel Reads Routing From a Frozen LLM Sep 15, 2026
  • AI math systems are still optimizing for the wrong proof Sep 12, 2026
  • Conversational XAI works best when it is a control surface, not a chatbot Sep 11, 2026
  • CRISPR screens need learned experiment pickers, not bigger chatbots Sep 11, 2026
  • DataShifts turns distribution shift into an error-budget problem Sep 11, 2026
  • MindTopo Tests Whether Vision Models Grasp Topology, Not Just Distances Sep 11, 2026
  • Recursive Self-Improvement Has a Roadmap Now, and Most of It Isn't Built Sep 11, 2026
  • ConvMem Turns Long-Context Reasoning Into a Tree Instead of a Chain Sep 10, 2026
  • Geometry reasoning gets better when the model is not doing every job Sep 10, 2026
  • OpenAI training-data accusations need a provenance test Sep 10, 2026
  • SG-JEPA points world models at the encoder, not the simulator Sep 10, 2026
  • NOAH models patient records as timelines, not snapshots Sep 9, 2026
  • Procedural Graphs give agents a memory of what to do, not just what happened Sep 9, 2026
  • AlphaGenome Atlas turns genome variants into an AI lookup layer Sep 8, 2026
  • OpenAI Says Its Own Researchers Now Lean on Coding Agents. What Does the Data Actually Show? Sep 6, 2026
  • Claude, Fermat, and the real value of machine-checkable proofs Sep 5, 2026
  • Auxiliary views explain why diverse pre-training data works Sep 4, 2026
  • Verbal reinforcement learning is a feedback routing problem Sep 2, 2026
  • SWE-Prime argues coding agents need cleaner wins, not more wins Aug 28, 2026
  • WikiSkill and the Case for Giving Agents a Memory That Compounds Aug 28, 2026
  • Anatomy-informed networks put clinical constraints before more data Aug 24, 2026
  • Agent memory works better when skills are small and written in text Aug 21, 2026
  • AlphaEvolve nudges the matrix multiplication exponent lower Aug 18, 2026
  • AutoSR turns symbolic regression into a research-state search Aug 18, 2026
  • PIHF turns prompts into a versioned policy layer Aug 18, 2026
  • YOPO makes abstention cheaper by reading the model before it lies Aug 17, 2026
  • Alignment baked into pretraining, not bolted on later Aug 14, 2026
  • AutoDesign turns paper-to-poster into a harness the agent rewrites itself Aug 14, 2026
  • LLMs Know When to Back Off, But Still Guess Too Specifically Aug 14, 2026
  • OmniScientist argues that AI scientists need eyes, not just workflows Aug 14, 2026
  • CLAUDE.md bloat is a memory problem, not a prompt problem Aug 12, 2026
  • Concise answers can weaken reasoning in fused LLM training Aug 11, 2026
  • Two prepared policies may be the sweet spot for uncertain MDPs Aug 4, 2026
  • AI research agents can code, but they still can’t judge the work Jul 30, 2026
  • The Automation Ceiling Nobody Prices In: When Human Participation Is the Product Jul 24, 2026
  • No Best Harness: Automated Discovery Systems Fail to Generalize Jul 21, 2026
  • Uncertainty metrics should follow the loss, not the other way around Jul 17, 2026
  • Vision models are getting the scene right, but not always the gaze Jul 13, 2026
  • Leanstral 1.5 points at proofs as a workflow, not a stunt Jul 11, 2026
  • DiaLLM separates dialect understanding from dialect writing Jul 9, 2026
  • Language critiques are a better training signal than a score, if you can afford them Jul 2, 2026
  • LLMs brainstorm like synthesis machines, not researchers Jul 2, 2026
  • TraceLab shows coding agents are an infrastructure workload now Jun 30, 2026
  • Google's Paper Assistant Wants to Catch Your Math Errors Before a Reviewer Does Jun 29, 2026
  • Nash solvers have preferences when the value is identical Jun 29, 2026
  • AI math is less about genius than verification Jun 27, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public