benchmarks
62 posts tagged benchmarks.
- What LLMs Miss About Haitian Creole, and Why Low-Resource Culture Breaks Evals
- When Documentation Doesn't Help Coding Agents: A Negative Result Worth Reading
- Swift Qwen’s speed claim is a local AI reminder, measure the whole loop
- ExplorationBench and the Gap Between Recall and Real Discovery
- The Modality Gap in Speech Fact-Checking, and Why Retrieval Alone Doesn't Fix It
- SWE-Flux: The Benchmark That Asks If Coding Models Can Predict What Code Actually Does
- CliffCompaction makes long coding runs cheaper by refusing to summarize
- Speaker-Centered Memory: Why Group Chats Break Your AI Agent
- DolphinBench Tests Agent Memory by Task, Not Trivia
- Where Harness Self-Improvement Actually Stands in Late 2026
- Full-Duplex Voice With Tool Calls: What NemotronLabs VoiceChat Actually Ships
- Benchmarks Do Not Settle the Open-Weight Risk Debate
- Decontamination reports are a promise, not a proof: the case for evaluation-side benchmarks
- Coding agents overclaim when their work is incomplete
- Paint-Anything makes hex colors a first-class diffusion control
- Double descent as an implicit regularization story
- MUSE tests vision-language models where classroom context gets messy
- When Rewording the Answer Key Reshuffles the Leaderboard
- Bellman Policy Optimization cuts one moving part from RLVR
- LLM personas fail when opinions have to change
- Real-SWE and the Case for Testing Coding Agents on Code They Have Never Seen
- An 86% Bitcoin quantum benchmark cut is not an 86% attack
- CRISPR screens need learned experiment pickers, not bigger chatbots
- MindTopo Tests Whether Vision Models Grasp Topology, Not Just Distances
- Geometry reasoning gets better when the model is not doing every job
- Instruction vs. Example: How Vision-Language Models Actually Moderate Content
- The Benchmark Says GPT-5. Your Users Get a Serving Route.
- Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet
- Terminal-Universe turns code-agent logs into reusable sandboxes
- Grading Coding Agents on How They Solve, Not Just Whether They Pass
- Terminal-Bench 4.0 and the eval gap for smaller coding agents
- CorporateBench and the Gap Between Demo Scale and Company Scale
- MCR-Bench shows code review agents still lose the plot
- Coding agents still struggle with whole-repo migrations
- Names Alone Break Entity Matching
- Prime Agent treats the harness as part of the model
- VIALS exposes the lab gap in vision-language models
- Qwen 3.8’s benchmark buzz is about reasoning budget, not vibes
- What AI4AI-Bench Says About Recursive Self-Improvement Right Now
- GLM-5.3 and the problem with “top open-weight coding model” claims
- VICBench shows vulnerability detection still needs humans
- GeoBenchLLM Tests Whether LLMs Understand Place
- Benchmarks Need QA Before They Judge Agents
- The Harness Is the Product: What HarnessOpt-Bench Actually Measures
- Video deep research agents need to look before they search
- DungeonBench puts tactical reasoning where agents usually break
- The Blind Spot in AI-Text Detectors: Human Writing an LLM Touched
- What FriendBench Reveals About How Models Read Social Cues
- Learning Seiberg dualities tests AI search, not physics vibes
- ORCA-bench shows oncall agents are not ready for pager duty
- APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups
- GPT-5.6’s ARC gain came from two API settings
- The Office-Task Benchmark That Prices Agents Against Human Labor
- Inkling’s first real test is not the headline benchmark
- Goal prompting is not a solver for NP-hard search
- Kimi K3’s SpreadsheetBench win is a signal, not a coronation
- The One-Shot Trap in Agent Optimization
- Voice AI needs a human-quality test, not another pretty demo
- Speech agents need timing tests, not just cleaner audio
- Buccy Bench makes model progress look usefully stupid
- Speaker recognition is a better agent test than another chat demo
- TestEvo-Bench moves coding-agent evals closer to real maintenance