Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

ai-safety

34 posts tagged ai-safety.

  • Layer-selective unlearning aims at the parts of a model that remember Sep 10, 2026
  • Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet Sep 9, 2026
  • AI safety should refuse harmful tasks, not entire topics Sep 8, 2026
  • Pachocki’s warning points to safety gates, not slower vibes Sep 8, 2026
  • OpenAI’s alignment note is really about operational discipline Sep 7, 2026
  • Uncensored Qwen edits show why model cards are not enough Sep 6, 2026
  • The OpenAI agent hack report is a disclosure problem first Sep 4, 2026
  • Bare assertions are a weak spot for medical reasoning models Sep 3, 2026
  • BLOOM-WILT and the Case for Auditing Models the Way Users Actually Break Them Sep 1, 2026
  • Certified world models still have blind topology Aug 31, 2026
  • AI influence ops are learning to fake institutions, not just posts Aug 25, 2026
  • Can Generated Text Prove Which Internal Path a Model Took? Aug 18, 2026
  • Model hypnosis turns harmless prompt quirks into a control surface Aug 18, 2026
  • RCI turns stop signals into safer offline RL training data Aug 13, 2026
  • Consistency checks are not truth checks for probabilistic AI Aug 12, 2026
  • What Six Years of TrustNLP Papers Say About Where AI Safety Research Actually Went Aug 12, 2026
  • Diffusion LLMs inherit the same brittle safety circuits Aug 10, 2026
  • OpenAI’s reported Astra pause is a cyber agent warning Aug 10, 2026
  • AI psychosis belongs in the workplace AI risk register Aug 8, 2026
  • OpenAI's cyber capability warning: what Astra's evals actually say Aug 8, 2026
  • OpenAI’s cyber eval issue is really a process story Aug 5, 2026
  • When Should a Robot Overrule Its Own Plan? CoWAM's Answer Aug 4, 2026
  • EPC scores explanations by testing what the model can lose Aug 3, 2026
  • Google Earth’s Nano Banana problem is provenance, not image quality Aug 1, 2026
  • What OpenAI's Cambodia Scam Takedown Tells Builders About Abuse Detection Aug 1, 2026
  • System prompts are becoming an audit surface Jul 31, 2026
  • The Same Model Name Gave Two Different Answers About Pseudo-Science Jul 27, 2026
  • Safety bounds are becoming probabilities, not vibes Jul 23, 2026
  • The Reconstruction Test Grades Vibes, Not Facts: Reading the RECAP Paper on Activation Explanations Jul 23, 2026
  • When a Model Eval Turns Into an Actual Breach Jul 22, 2026
  • LLM agents may change their answers when the room changes Jul 3, 2026
  • The Safety Case for a Boring Threshold Jul 3, 2026
  • Safety evals need to test messy language, not just bad behavior Jul 2, 2026
  • Synthetic QA Has a Selection Problem Before It Has a Training Problem Jul 1, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public