Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

interpretability

26 posts tagged interpretability.

  • Flow matching gives neural dissimilarity metrics one common frame Sep 28, 2026
  • Models may refuse based on who they think you are Sep 28, 2026
  • Audio models do not automatically share speech and text concepts Sep 25, 2026
  • When a model’s rejection reason changes its next choice Sep 25, 2026
  • Clifford-VAE puts pixels into symbolic memory space Sep 24, 2026
  • When a Model Gets Math Right but Reads the Same Problem Differently Sep 24, 2026
  • DiaVLo Turns Vision-Language Model Failures Into Named Behaviours Sep 21, 2026
  • Membership Inference Gets an Entropy Correction: Reading the ETD Paper Sep 21, 2026
  • Probabilistic Linear Explanations make interpretability more honest Sep 17, 2026
  • Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet Sep 9, 2026
  • CAST makes clinical model audits more concrete Aug 28, 2026
  • What a Moral Probe Finds Inside an LLM: Structure, Not a Single Switch Aug 28, 2026
  • Can Generated Text Prove Which Internal Path a Model Took? Aug 18, 2026
  • Model hypnosis turns harmless prompt quirks into a control surface Aug 18, 2026
  • What Six Years of TrustNLP Papers Say About Where AI Safety Research Actually Went Aug 12, 2026
  • CENDRe Brings Frequency-Domain Explanations to Time-Series CNNs Aug 3, 2026
  • The Gap Between Reading a Feature and Steering With It Jul 28, 2026
  • The Reconstruction Test Grades Vibes, Not Facts: Reading the RECAP Paper on Activation Explanations Jul 23, 2026
  • Judge Bias Lives in the Activations, Not Just the Prompt Jul 14, 2026
  • Transformer circuits may be lower-dimensional than they look Jul 14, 2026
  • Language model embeddings do not want to collapse Jul 13, 2026
  • SciReasoner Treats Molecular Structure as Evidence You Can Inspect Jul 9, 2026
  • C2R targets the hidden mess inside sparse autoencoder features Jun 30, 2026
  • Reasoning Traces as a Difficulty Sensor: What Epi2Diff Gets Right Jun 29, 2026
  • When Models Quietly Unlearn: The Natural Ungrokking Problem Jun 25, 2026
  • Reading Gradients to Catch Hallucinations Before They Ship Jun 24, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public