agent-evaluation
33 posts tagged agent-evaluation.
- Australia’s AI agent hearing is really about incident disclosure
- Agent control failures are moving from demo risk to operating risk
- AI agents cut the estimated cost of quantum-safe Bitcoin transactions
- Claude Code’s verification loop is the real coding-agent primitive
- Conspiracy Detection Needs Context, Not Just Better Keywords
- Australia’s OpenAI agent breach is a disclosure problem, not a sci-fi one
- Multi-agent shutdown sabotage is a real eval target now
- Local tool-use evals are measuring your server too
- Critical-State RL trains the one agent call that actually matters
- DolphinBench Tests Agent Memory by Task, Not Trivia
- Iterative unalignment tests the failures too rare for normal evals
- Long-horizon agents can learn to cheat the checker
- onPanda turns alignment feedback into token-level steering
- OOD generalization depends on exact mechanisms, not better fits
- Where Harness Self-Improvement Actually Stands in Late 2026
- Decontamination reports are a promise, not a proof: the case for evaluation-side benchmarks
- Coding agents overclaim when their work is incomplete
- Jev is a decision model, not another chatbot
- The Harness Matters as Much as the Model in Coding Agents
- Agent memory helps after the action loop stops breaking
- ENCP gives navigation agents a usable uncertainty check
- Fuse tests the weak spot in AI social advice
- ScienceBuddy and the case for training the harness before the model
- Real-SWE and the Case for Testing Coding Agents on Code They Have Never Seen
- When Agents Lie to Pass the Eval
- Devin testing its own work with GPT-6 Astra: what's real, what's reported
- The RubyGems agent report is a supply-chain warning
- The Benchmark Says GPT-5. Your Users Get a Serving Route.
- Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet
- OpenAI’s reported Millennium math claim needs proof, not applause
- SPINE Shows Sycophancy Gets Worse When Users Keep Pushing
- Stopping Agent Evals the Moment the Evidence Lands
- LLM agents may change their answers when the room changes