Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

agent-evaluation

33 posts tagged agent-evaluation.

  • Australia’s AI agent hearing is really about incident disclosure Sep 28, 2026
  • Agent control failures are moving from demo risk to operating risk Sep 27, 2026
  • AI agents cut the estimated cost of quantum-safe Bitcoin transactions Sep 26, 2026
  • Claude Code’s verification loop is the real coding-agent primitive Sep 26, 2026
  • Conspiracy Detection Needs Context, Not Just Better Keywords Sep 25, 2026
  • Australia’s OpenAI agent breach is a disclosure problem, not a sci-fi one Sep 24, 2026
  • Multi-agent shutdown sabotage is a real eval target now Sep 24, 2026
  • Local tool-use evals are measuring your server too Sep 23, 2026
  • Critical-State RL trains the one agent call that actually matters Sep 22, 2026
  • DolphinBench Tests Agent Memory by Task, Not Trivia Sep 22, 2026
  • Iterative unalignment tests the failures too rare for normal evals Sep 22, 2026
  • Long-horizon agents can learn to cheat the checker Sep 22, 2026
  • onPanda turns alignment feedback into token-level steering Sep 22, 2026
  • OOD generalization depends on exact mechanisms, not better fits Sep 22, 2026
  • Where Harness Self-Improvement Actually Stands in Late 2026 Sep 22, 2026
  • Decontamination reports are a promise, not a proof: the case for evaluation-side benchmarks Sep 20, 2026
  • Coding agents overclaim when their work is incomplete Sep 18, 2026
  • Jev is a decision model, not another chatbot Sep 18, 2026
  • The Harness Matters as Much as the Model in Coding Agents Sep 18, 2026
  • Agent memory helps after the action loop stops breaking Sep 17, 2026
  • ENCP gives navigation agents a usable uncertainty check Sep 16, 2026
  • Fuse tests the weak spot in AI social advice Sep 16, 2026
  • ScienceBuddy and the case for training the harness before the model Sep 16, 2026
  • Real-SWE and the Case for Testing Coding Agents on Code They Have Never Seen Sep 13, 2026
  • When Agents Lie to Pass the Eval Sep 13, 2026
  • Devin testing its own work with GPT-6 Astra: what's real, what's reported Sep 12, 2026
  • The RubyGems agent report is a supply-chain warning Sep 12, 2026
  • The Benchmark Says GPT-5. Your Users Get a Serving Route. Sep 10, 2026
  • Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet Sep 9, 2026
  • OpenAI’s reported Millennium math claim needs proof, not applause Sep 9, 2026
  • SPINE Shows Sycophancy Gets Worse When Users Keep Pushing Sep 9, 2026
  • Stopping Agent Evals the Moment the Evidence Lands Aug 7, 2026
  • LLM agents may change their answers when the room changes Jul 3, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public