Skip to content
{ ken ashe }
  • Building
  • Topics
  • Blog
  • Building
  • Topics
  • Blog
← Blog / Tags

// tag

ai-safety

21 posts tagged ai-safety.

  • RCI turns stop signals into safer offline RL training data Aug 13, 2026
  • Consistency checks are not truth checks for probabilistic AI Aug 12, 2026
  • What Six Years of TrustNLP Papers Say About Where AI Safety Research Actually Went Aug 12, 2026
  • Diffusion LLMs inherit the same brittle safety circuits Aug 10, 2026
  • OpenAI’s reported Astra pause is a cyber agent warning Aug 10, 2026
  • AI psychosis belongs in the workplace AI risk register Aug 8, 2026
  • OpenAI's cyber capability warning: what Astra's evals actually say Aug 8, 2026
  • OpenAI’s cyber eval issue is really a process story Aug 5, 2026
  • When Should a Robot Overrule Its Own Plan? CoWAM's Answer Aug 4, 2026
  • EPC scores explanations by testing what the model can lose Aug 3, 2026
  • Google Earth’s Nano Banana problem is provenance, not image quality Aug 1, 2026
  • What OpenAI's Cambodia Scam Takedown Tells Builders About Abuse Detection Aug 1, 2026
  • System prompts are becoming an audit surface Jul 31, 2026
  • The Same Model Name Gave Two Different Answers About Pseudo-Science Jul 27, 2026
  • Safety bounds are becoming probabilities, not vibes Jul 23, 2026
  • The Reconstruction Test Grades Vibes, Not Facts: Reading the RECAP Paper on Activation Explanations Jul 23, 2026
  • When a Model Eval Turns Into an Actual Breach Jul 22, 2026
  • LLM agents may change their answers when the room changes Jul 3, 2026
  • The Safety Case for a Boring Threshold Jul 3, 2026
  • Safety evals need to test messy language, not just bad behavior Jul 2, 2026
  • Synthetic QA Has a Selection Problem Before It Has a Training Problem Jul 1, 2026
  • RSS
  • LinkedIn
  • X
  • GitHub
  • Email
  • Newsroom
  • Media Kit
  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe