benchmarks
22 posts tagged benchmarks.
- VICBench shows vulnerability detection still needs humans
- GeoBenchLLM Tests Whether LLMs Understand Place
- Benchmarks Need QA Before They Judge Agents
- The Harness Is the Product: What HarnessOpt-Bench Actually Measures
- Video deep research agents need to look before they search
- DungeonBench puts tactical reasoning where agents usually break
- The Blind Spot in AI-Text Detectors: Human Writing an LLM Touched
- What FriendBench Reveals About How Models Read Social Cues
- Learning Seiberg dualities tests AI search, not physics vibes
- ORCA-bench shows oncall agents are not ready for pager duty
- APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups
- GPT-5.6’s ARC gain came from two API settings
- The Office-Task Benchmark That Prices Agents Against Human Labor
- Inkling’s first real test is not the headline benchmark
- Goal prompting is not a solver for NP-hard search
- Kimi K3’s SpreadsheetBench win is a signal, not a coronation
- The One-Shot Trap in Agent Optimization
- Voice AI needs a human-quality test, not another pretty demo
- Speech agents need timing tests, not just cleaner audio
- Buccy Bench makes model progress look usefully stupid
- Speaker recognition is a better agent test than another chat demo
- TestEvo-Bench moves coding-agent evals closer to real maintenance