Skip to content
{ ken ashe }
  • Building
  • Topics
  • Blog
  • Building
  • Topics
  • Blog
← Blog / Tags

// tag

evaluation

18 posts tagged evaluation.

  • Auto-research agents need fuzzer-style feedback Aug 11, 2026
  • The TTS Score That Hides Your Voice Model's Real Problems Aug 11, 2026
  • TutorMoments and the Hard Part of AI Tutoring: Knowing When Not to Answer Aug 8, 2026
  • Agnostic PAC learning gets its optimal bound Aug 7, 2026
  • Selective Trust: Why RAG Systems Should Learn When to Ignore Their Own Context Aug 7, 2026
  • What 'Test-Time Scaling' Actually Means When You Read a Benchmark Aug 5, 2026
  • Computer-use agents need stricter judges, not prettier demos Jul 31, 2026
  • AI research agents can code, but they still can’t judge the work Jul 30, 2026
  • Autonomous research needs a budget scoreboard Jul 28, 2026
  • No Best Harness: Automated Discovery Systems Fail to Generalize Jul 21, 2026
  • Memory Failures Hide Behind Correct Answers: What MemOps Exposes Jul 15, 2026
  • Self-repair in small code models may be measuring retry form, not error content Jul 15, 2026
  • Voice AI needs a human-quality test, not another pretty demo Jul 15, 2026
  • Judge Bias Lives in the Activations, Not Just the Prompt Jul 14, 2026
  • Vision models are getting the scene right, but not always the gaze Jul 13, 2026
  • DiaLLM separates dialect understanding from dialect writing Jul 9, 2026
  • Unlearning That Hides Instead of Erases: What LACUNA Exposes Jul 3, 2026
  • The Cheap Way to Test Whether Your Agent's Step Scores Actually Mean Anything Jul 1, 2026
  • RSS
  • LinkedIn
  • X
  • GitHub
  • Email
  • Newsroom
  • Media Kit
  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe