interpretability
12 posts tagged interpretability.
- What Six Years of TrustNLP Papers Say About Where AI Safety Research Actually Went
- CENDRe Brings Frequency-Domain Explanations to Time-Series CNNs
- The Gap Between Reading a Feature and Steering With It
- The Reconstruction Test Grades Vibes, Not Facts: Reading the RECAP Paper on Activation Explanations
- Judge Bias Lives in the Activations, Not Just the Prompt
- Transformer circuits may be lower-dimensional than they look
- Language model embeddings do not want to collapse
- SciReasoner Treats Molecular Structure as Evidence You Can Inspect
- C2R targets the hidden mess inside sparse autoencoder features
- Reasoning Traces as a Difficulty Sensor: What Epi2Diff Gets Right
- When Models Quietly Unlearn: The Natural Ungrokking Problem
- Reading Gradients to Catch Hallucinations Before They Ship