interpretability
26 posts tagged interpretability.
- Flow matching gives neural dissimilarity metrics one common frame
- Models may refuse based on who they think you are
- Audio models do not automatically share speech and text concepts
- When a model’s rejection reason changes its next choice
- Clifford-VAE puts pixels into symbolic memory space
- When a Model Gets Math Right but Reads the Same Problem Differently
- DiaVLo Turns Vision-Language Model Failures Into Named Behaviours
- Membership Inference Gets an Entropy Correction: Reading the ETD Paper
- Probabilistic Linear Explanations make interpretability more honest
- Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet
- CAST makes clinical model audits more concrete
- What a Moral Probe Finds Inside an LLM: Structure, Not a Single Switch
- Can Generated Text Prove Which Internal Path a Model Took?
- Model hypnosis turns harmless prompt quirks into a control surface
- What Six Years of TrustNLP Papers Say About Where AI Safety Research Actually Went
- CENDRe Brings Frequency-Domain Explanations to Time-Series CNNs
- The Gap Between Reading a Feature and Steering With It
- The Reconstruction Test Grades Vibes, Not Facts: Reading the RECAP Paper on Activation Explanations
- Judge Bias Lives in the Activations, Not Just the Prompt
- Transformer circuits may be lower-dimensional than they look
- Language model embeddings do not want to collapse
- SciReasoner Treats Molecular Structure as Evidence You Can Inspect
- C2R targets the hidden mess inside sparse autoencoder features
- Reasoning Traces as a Difficulty Sensor: What Epi2Diff Gets Right
- When Models Quietly Unlearn: The Natural Ungrokking Problem
- Reading Gradients to Catch Hallucinations Before They Ship