evaluation
18 posts tagged evaluation.
- Auto-research agents need fuzzer-style feedback
- The TTS Score That Hides Your Voice Model's Real Problems
- TutorMoments and the Hard Part of AI Tutoring: Knowing When Not to Answer
- Agnostic PAC learning gets its optimal bound
- Selective Trust: Why RAG Systems Should Learn When to Ignore Their Own Context
- What 'Test-Time Scaling' Actually Means When You Read a Benchmark
- Computer-use agents need stricter judges, not prettier demos
- AI research agents can code, but they still can’t judge the work
- Autonomous research needs a budget scoreboard
- No Best Harness: Automated Discovery Systems Fail to Generalize
- Memory Failures Hide Behind Correct Answers: What MemOps Exposes
- Self-repair in small code models may be measuring retry form, not error content
- Voice AI needs a human-quality test, not another pretty demo
- Judge Bias Lives in the Activations, Not Just the Prompt
- Vision models are getting the scene right, but not always the gaze
- DiaLLM separates dialect understanding from dialect writing
- Unlearning That Hides Instead of Erases: What LACUNA Exposes
- The Cheap Way to Test Whether Your Agent's Step Scores Actually Mean Anything