Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

evaluation

28 posts tagged evaluation.

  • LexFlip Shows Legal Meaning Metrics Are Still Token Counters Sep 7, 2026
  • AICOME works best when the missing variable is narrow Sep 3, 2026
  • When a Language Model's Odds Don't Add Up Sep 3, 2026
  • Grading Coding Agents on How They Solve, Not Just Whether They Pass Sep 2, 2026
  • Reward choice changes how LLM forecasters are wrong Aug 31, 2026
  • Synthetic data needs a weight limit Aug 31, 2026
  • A Minecraft clone is a weak coding-agent test Aug 30, 2026
  • The Self-Improving Agent Demo That Falls Apart When You Shuffle the Tasks Aug 19, 2026
  • CaliBench tests whether video models roll fair dice Aug 18, 2026
  • Stop Making One Prompt Read the Sources and Cast the Vote Aug 17, 2026
  • Auto-research agents need fuzzer-style feedback Aug 11, 2026
  • The TTS Score That Hides Your Voice Model's Real Problems Aug 11, 2026
  • TutorMoments and the Hard Part of AI Tutoring: Knowing When Not to Answer Aug 8, 2026
  • Agnostic PAC learning gets its optimal bound Aug 7, 2026
  • Selective Trust: Why RAG Systems Should Learn When to Ignore Their Own Context Aug 7, 2026
  • What 'Test-Time Scaling' Actually Means When You Read a Benchmark Aug 5, 2026
  • Computer-use agents need stricter judges, not prettier demos Jul 31, 2026
  • AI research agents can code, but they still can’t judge the work Jul 30, 2026
  • Autonomous research needs a budget scoreboard Jul 28, 2026
  • No Best Harness: Automated Discovery Systems Fail to Generalize Jul 21, 2026
  • Memory Failures Hide Behind Correct Answers: What MemOps Exposes Jul 15, 2026
  • Self-repair in small code models may be measuring retry form, not error content Jul 15, 2026
  • Voice AI needs a human-quality test, not another pretty demo Jul 15, 2026
  • Judge Bias Lives in the Activations, Not Just the Prompt Jul 14, 2026
  • Vision models are getting the scene right, but not always the gaze Jul 13, 2026
  • DiaLLM separates dialect understanding from dialect writing Jul 9, 2026
  • Unlearning That Hides Instead of Erases: What LACUNA Exposes Jul 3, 2026
  • The Cheap Way to Test Whether Your Agent's Step Scores Actually Mean Anything Jul 1, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public