ai-safety
34 posts tagged ai-safety.
- Layer-selective unlearning aims at the parts of a model that remember
- Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet
- AI safety should refuse harmful tasks, not entire topics
- Pachocki’s warning points to safety gates, not slower vibes
- OpenAI’s alignment note is really about operational discipline
- Uncensored Qwen edits show why model cards are not enough
- The OpenAI agent hack report is a disclosure problem first
- Bare assertions are a weak spot for medical reasoning models
- BLOOM-WILT and the Case for Auditing Models the Way Users Actually Break Them
- Certified world models still have blind topology
- AI influence ops are learning to fake institutions, not just posts
- Can Generated Text Prove Which Internal Path a Model Took?
- Model hypnosis turns harmless prompt quirks into a control surface
- RCI turns stop signals into safer offline RL training data
- Consistency checks are not truth checks for probabilistic AI
- What Six Years of TrustNLP Papers Say About Where AI Safety Research Actually Went
- Diffusion LLMs inherit the same brittle safety circuits
- OpenAI’s reported Astra pause is a cyber agent warning
- AI psychosis belongs in the workplace AI risk register
- OpenAI's cyber capability warning: what Astra's evals actually say
- OpenAI’s cyber eval issue is really a process story
- When Should a Robot Overrule Its Own Plan? CoWAM's Answer
- EPC scores explanations by testing what the model can lose
- Google Earth’s Nano Banana problem is provenance, not image quality
- What OpenAI's Cambodia Scam Takedown Tells Builders About Abuse Detection
- System prompts are becoming an audit surface
- The Same Model Name Gave Two Different Answers About Pseudo-Science
- Safety bounds are becoming probabilities, not vibes
- The Reconstruction Test Grades Vibes, Not Facts: Reading the RECAP Paper on Activation Explanations
- When a Model Eval Turns Into an Actual Breach
- LLM agents may change their answers when the room changes
- The Safety Case for a Boring Threshold
- Safety evals need to test messy language, not just bad behavior
- Synthetic QA Has a Selection Problem Before It Has a Training Problem