Skip to content
{ ken ashe }
  • Building
  • Topics
  • Blog
  • Building
  • Topics
  • Blog
← Blog / Tags

// tag

benchmarks

22 posts tagged benchmarks.

  • VICBench shows vulnerability detection still needs humans Aug 13, 2026
  • GeoBenchLLM Tests Whether LLMs Understand Place Aug 10, 2026
  • Benchmarks Need QA Before They Judge Agents Aug 7, 2026
  • The Harness Is the Product: What HarnessOpt-Bench Actually Measures Aug 7, 2026
  • Video deep research agents need to look before they search Aug 5, 2026
  • DungeonBench puts tactical reasoning where agents usually break Aug 3, 2026
  • The Blind Spot in AI-Text Detectors: Human Writing an LLM Touched Aug 3, 2026
  • What FriendBench Reveals About How Models Read Social Cues Aug 3, 2026
  • Learning Seiberg dualities tests AI search, not physics vibes Jul 31, 2026
  • ORCA-bench shows oncall agents are not ready for pager duty Jul 31, 2026
  • APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups Jul 30, 2026
  • GPT-5.6’s ARC gain came from two API settings Jul 30, 2026
  • The Office-Task Benchmark That Prices Agents Against Human Labor Jul 30, 2026
  • Inkling’s first real test is not the headline benchmark Jul 26, 2026
  • Goal prompting is not a solver for NP-hard search Jul 19, 2026
  • Kimi K3’s SpreadsheetBench win is a signal, not a coronation Jul 19, 2026
  • The One-Shot Trap in Agent Optimization Jul 16, 2026
  • Voice AI needs a human-quality test, not another pretty demo Jul 15, 2026
  • Speech agents need timing tests, not just cleaner audio Jul 7, 2026
  • Buccy Bench makes model progress look usefully stupid Jul 4, 2026
  • Speaker recognition is a better agent test than another chat demo Jul 3, 2026
  • TestEvo-Bench moves coding-agent evals closer to real maintenance Jul 3, 2026
  • RSS
  • LinkedIn
  • X
  • GitHub
  • Email
  • Newsroom
  • Media Kit
  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe