ai-safety
21 posts tagged ai-safety.
- RCI turns stop signals into safer offline RL training data
- Consistency checks are not truth checks for probabilistic AI
- What Six Years of TrustNLP Papers Say About Where AI Safety Research Actually Went
- Diffusion LLMs inherit the same brittle safety circuits
- OpenAI’s reported Astra pause is a cyber agent warning
- AI psychosis belongs in the workplace AI risk register
- OpenAI's cyber capability warning: what Astra's evals actually say
- OpenAI’s cyber eval issue is really a process story
- When Should a Robot Overrule Its Own Plan? CoWAM's Answer
- EPC scores explanations by testing what the model can lose
- Google Earth’s Nano Banana problem is provenance, not image quality
- What OpenAI's Cambodia Scam Takedown Tells Builders About Abuse Detection
- System prompts are becoming an audit surface
- The Same Model Name Gave Two Different Answers About Pseudo-Science
- Safety bounds are becoming probabilities, not vibes
- The Reconstruction Test Grades Vibes, Not Facts: Reading the RECAP Paper on Activation Explanations
- When a Model Eval Turns Into an Actual Breach
- LLM agents may change their answers when the room changes
- The Safety Case for a Boring Threshold
- Safety evals need to test messy language, not just bad behavior
- Synthetic QA Has a Selection Problem Before It Has a Training Problem