Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

ai-safety

52 posts tagged ai-safety.

  • Black-box attribute alignment is a sampler, not a fairness wand Sep 28, 2026
  • Models may refuse based on who they think you are Sep 28, 2026
  • Multi-agent shutdown sabotage is a real eval target now Sep 24, 2026
  • Iterative unalignment tests the failures too rare for normal evals Sep 22, 2026
  • Long-horizon agents can learn to cheat the checker Sep 22, 2026
  • OpenAI wants shared AI standards, but the hard part is enforcement Sep 21, 2026
  • Benchmarks Do Not Settle the Open-Weight Risk Debate Sep 20, 2026
  • OpenAI's Australian Youth Safety Blueprint: what six pillars mean for anyone building AI products Sep 19, 2026
  • Toxicity Scores Can Miss Sanitized Bias in GPT Outputs Sep 18, 2026
  • When a Coding Agent Drives a Robot, Task Success Isn't Safety Sep 18, 2026
  • Teaching Models to Say 'I Don't Know' With a Prompt, Not a Retrain Sep 16, 2026
  • Zuckerberg’s case for solo AI slowdowns Sep 16, 2026
  • AI slowdown promises break without enforcement Sep 15, 2026
  • K-Bench tests mental health chatbots where generic safety evals do not Sep 15, 2026
  • The Thin Claim Behind “No Regulation” for Frontier Model Pacing Sep 14, 2026
  • Trump’s AI Guardrails Argument Turns Safety Into a Power Question Sep 14, 2026
  • AI leaders are asking for a slower race before recursive systems arrive Sep 13, 2026
  • The Only Safe Pace for AI Is Everyone Else's Sep 13, 2026
  • Layer-selective unlearning aims at the parts of a model that remember Sep 10, 2026
  • Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet Sep 9, 2026
  • AI safety should refuse harmful tasks, not entire topics Sep 8, 2026
  • Pachocki’s warning points to safety gates, not slower vibes Sep 8, 2026
  • OpenAI’s alignment note is really about operational discipline Sep 7, 2026
  • Uncensored Qwen edits show why model cards are not enough Sep 6, 2026
  • The OpenAI agent hack report is a disclosure problem first Sep 4, 2026
  • Bare assertions are a weak spot for medical reasoning models Sep 3, 2026
  • BLOOM-WILT and the Case for Auditing Models the Way Users Actually Break Them Sep 1, 2026
  • Certified world models still have blind topology Aug 31, 2026
  • AI influence ops are learning to fake institutions, not just posts Aug 25, 2026
  • Can Generated Text Prove Which Internal Path a Model Took? Aug 18, 2026
  • Model hypnosis turns harmless prompt quirks into a control surface Aug 18, 2026
  • RCI turns stop signals into safer offline RL training data Aug 13, 2026
  • Consistency checks are not truth checks for probabilistic AI Aug 12, 2026
  • What Six Years of TrustNLP Papers Say About Where AI Safety Research Actually Went Aug 12, 2026
  • Diffusion LLMs inherit the same brittle safety circuits Aug 10, 2026
  • OpenAI’s reported Astra pause is a cyber agent warning Aug 10, 2026
  • AI psychosis belongs in the workplace AI risk register Aug 8, 2026
  • OpenAI's cyber capability warning: what Astra's evals actually say Aug 8, 2026
  • OpenAI’s cyber eval issue is really a process story Aug 5, 2026
  • When Should a Robot Overrule Its Own Plan? CoWAM's Answer Aug 4, 2026
  • EPC scores explanations by testing what the model can lose Aug 3, 2026
  • Google Earth’s Nano Banana problem is provenance, not image quality Aug 1, 2026
  • What OpenAI's Cambodia Scam Takedown Tells Builders About Abuse Detection Aug 1, 2026
  • System prompts are becoming an audit surface Jul 31, 2026
  • The Same Model Name Gave Two Different Answers About Pseudo-Science Jul 27, 2026
  • Safety bounds are becoming probabilities, not vibes Jul 23, 2026
  • The Reconstruction Test Grades Vibes, Not Facts: Reading the RECAP Paper on Activation Explanations Jul 23, 2026
  • When a Model Eval Turns Into an Actual Breach Jul 22, 2026
  • LLM agents may change their answers when the room changes Jul 3, 2026
  • The Safety Case for a Boring Threshold Jul 3, 2026
  • Safety evals need to test messy language, not just bad behavior Jul 2, 2026
  • Synthetic QA Has a Selection Problem Before It Has a Training Problem Jul 1, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public