Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

benchmarks

62 posts tagged benchmarks.

  • What LLMs Miss About Haitian Creole, and Why Low-Resource Culture Breaks Evals Sep 28, 2026
  • When Documentation Doesn't Help Coding Agents: A Negative Result Worth Reading Sep 28, 2026
  • Swift Qwen’s speed claim is a local AI reminder, measure the whole loop Sep 26, 2026
  • ExplorationBench and the Gap Between Recall and Real Discovery Sep 25, 2026
  • The Modality Gap in Speech Fact-Checking, and Why Retrieval Alone Doesn't Fix It Sep 25, 2026
  • SWE-Flux: The Benchmark That Asks If Coding Models Can Predict What Code Actually Does Sep 24, 2026
  • CliffCompaction makes long coding runs cheaper by refusing to summarize Sep 23, 2026
  • Speaker-Centered Memory: Why Group Chats Break Your AI Agent Sep 23, 2026
  • DolphinBench Tests Agent Memory by Task, Not Trivia Sep 22, 2026
  • Where Harness Self-Improvement Actually Stands in Late 2026 Sep 22, 2026
  • Full-Duplex Voice With Tool Calls: What NemotronLabs VoiceChat Actually Ships Sep 21, 2026
  • Benchmarks Do Not Settle the Open-Weight Risk Debate Sep 20, 2026
  • Decontamination reports are a promise, not a proof: the case for evaluation-side benchmarks Sep 20, 2026
  • Coding agents overclaim when their work is incomplete Sep 18, 2026
  • Paint-Anything makes hex colors a first-class diffusion control Sep 18, 2026
  • Double descent as an implicit regularization story Sep 17, 2026
  • MUSE tests vision-language models where classroom context gets messy Sep 17, 2026
  • When Rewording the Answer Key Reshuffles the Leaderboard Sep 17, 2026
  • Bellman Policy Optimization cuts one moving part from RLVR Sep 15, 2026
  • LLM personas fail when opinions have to change Sep 15, 2026
  • Real-SWE and the Case for Testing Coding Agents on Code They Have Never Seen Sep 13, 2026
  • An 86% Bitcoin quantum benchmark cut is not an 86% attack Sep 11, 2026
  • CRISPR screens need learned experiment pickers, not bigger chatbots Sep 11, 2026
  • MindTopo Tests Whether Vision Models Grasp Topology, Not Just Distances Sep 11, 2026
  • Geometry reasoning gets better when the model is not doing every job Sep 10, 2026
  • Instruction vs. Example: How Vision-Language Models Actually Moderate Content Sep 10, 2026
  • The Benchmark Says GPT-5. Your Users Get a Serving Route. Sep 10, 2026
  • Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet Sep 9, 2026
  • Terminal-Universe turns code-agent logs into reusable sandboxes Sep 4, 2026
  • Grading Coding Agents on How They Solve, Not Just Whether They Pass Sep 2, 2026
  • Terminal-Bench 4.0 and the eval gap for smaller coding agents Aug 29, 2026
  • CorporateBench and the Gap Between Demo Scale and Company Scale Aug 28, 2026
  • MCR-Bench shows code review agents still lose the plot Aug 28, 2026
  • Coding agents still struggle with whole-repo migrations Aug 25, 2026
  • Names Alone Break Entity Matching Aug 25, 2026
  • Prime Agent treats the harness as part of the model Aug 25, 2026
  • VIALS exposes the lab gap in vision-language models Aug 24, 2026
  • Qwen 3.8’s benchmark buzz is about reasoning budget, not vibes Aug 22, 2026
  • What AI4AI-Bench Says About Recursive Self-Improvement Right Now Aug 21, 2026
  • GLM-5.3 and the problem with “top open-weight coding model” claims Aug 15, 2026
  • VICBench shows vulnerability detection still needs humans Aug 13, 2026
  • GeoBenchLLM Tests Whether LLMs Understand Place Aug 10, 2026
  • Benchmarks Need QA Before They Judge Agents Aug 7, 2026
  • The Harness Is the Product: What HarnessOpt-Bench Actually Measures Aug 7, 2026
  • Video deep research agents need to look before they search Aug 5, 2026
  • DungeonBench puts tactical reasoning where agents usually break Aug 3, 2026
  • The Blind Spot in AI-Text Detectors: Human Writing an LLM Touched Aug 3, 2026
  • What FriendBench Reveals About How Models Read Social Cues Aug 3, 2026
  • Learning Seiberg dualities tests AI search, not physics vibes Jul 31, 2026
  • ORCA-bench shows oncall agents are not ready for pager duty Jul 31, 2026
  • APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups Jul 30, 2026
  • GPT-5.6’s ARC gain came from two API settings Jul 30, 2026
  • The Office-Task Benchmark That Prices Agents Against Human Labor Jul 30, 2026
  • Inkling’s first real test is not the headline benchmark Jul 26, 2026
  • Goal prompting is not a solver for NP-hard search Jul 19, 2026
  • Kimi K3’s SpreadsheetBench win is a signal, not a coronation Jul 19, 2026
  • The One-Shot Trap in Agent Optimization Jul 16, 2026
  • Voice AI needs a human-quality test, not another pretty demo Jul 15, 2026
  • Speech agents need timing tests, not just cleaner audio Jul 7, 2026
  • Buccy Bench makes model progress look usefully stupid Jul 4, 2026
  • Speaker recognition is a better agent test than another chat demo Jul 3, 2026
  • TestEvo-Bench moves coding-agent evals closer to real maintenance Jul 3, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public