Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

evals

76 posts tagged evals.

  • Flow matching gives neural dissimilarity metrics one common frame Sep 28, 2026
  • What LLMs Miss About Haitian Creole, and Why Low-Resource Culture Breaks Evals Sep 28, 2026
  • NSA’s reported AI testing spend makes evals look like infrastructure Sep 26, 2026
  • ExplorationBench and the Gap Between Recall and Real Discovery Sep 25, 2026
  • When a model’s rejection reason changes its next choice Sep 25, 2026
  • SWE-Flux: The Benchmark That Asks If Coding Models Can Predict What Code Actually Does Sep 24, 2026
  • Translation Fine-Tuning Breaks the Controls General Evals Miss Sep 24, 2026
  • When a Model Gets Math Right but Reads the Same Problem Differently Sep 24, 2026
  • Long-context models can miss the answer sitting behind nearby noise Sep 23, 2026
  • DolphinBench Tests Agent Memory by Task, Not Trivia Sep 22, 2026
  • Iterative unalignment tests the failures too rare for normal evals Sep 22, 2026
  • Jev shows why science workflows need semantic evals, not just final-answer grading Sep 22, 2026
  • DiaVLo Turns Vision-Language Model Failures Into Named Behaviours Sep 21, 2026
  • Multi-hop RAG needs confidence before the answer Sep 21, 2026
  • OpenAI wants shared AI standards, but the hard part is enforcement Sep 21, 2026
  • OpenAI’s math advisory group is about claim control Sep 21, 2026
  • Decontamination reports are a promise, not a proof: the case for evaluation-side benchmarks Sep 20, 2026
  • A cipher win is not an eval without the working Sep 19, 2026
  • Toxicity Scores Can Miss Sanitized Bias in GPT Outputs Sep 18, 2026
  • Probabilistic Linear Explanations make interpretability more honest Sep 17, 2026
  • When Rewording the Answer Key Reshuffles the Leaderboard Sep 17, 2026
  • Teaching Models to Say 'I Don't Know' With a Prompt, Not a Retrain Sep 16, 2026
  • K-Bench tests mental health chatbots where generic safety evals do not Sep 15, 2026
  • LLM personas fail when opinions have to change Sep 15, 2026
  • When Agents Lie to Pass the Eval Sep 13, 2026
  • AI math systems are still optimizing for the wrong proof Sep 12, 2026
  • GPT-6 Astra and the problem with moving-target models Sep 12, 2026
  • DataShifts turns distribution shift into an error-budget problem Sep 11, 2026
  • GEO needs exposure data before it can claim revenue Sep 11, 2026
  • The Hallucination Detector That Doesn't Transfer to Your Domain Sep 11, 2026
  • Instruction vs. Example: How Vision-Language Models Actually Moderate Content Sep 10, 2026
  • OpenAI training-data accusations need a provenance test Sep 10, 2026
  • The Benchmark Says GPT-5. Your Users Get a Serving Route. Sep 10, 2026
  • SPINE Shows Sycophancy Gets Worse When Users Keep Pushing Sep 9, 2026
  • LexFlip Shows Legal Meaning Metrics Are Still Token Counters Sep 7, 2026
  • Hosted LLM judges are not fixed instruments Sep 4, 2026
  • AICOME works best when the missing variable is narrow Sep 3, 2026
  • Bare assertions are a weak spot for medical reasoning models Sep 3, 2026
  • When a Language Model's Odds Don't Add Up Sep 3, 2026
  • Grading Coding Agents on How They Solve, Not Just Whether They Pass Sep 2, 2026
  • Reward choice changes how LLM forecasters are wrong Aug 31, 2026
  • Synthetic data needs a weight limit Aug 31, 2026
  • A Minecraft clone is a weak coding-agent test Aug 30, 2026
  • Terminal-Bench 4.0 and the eval gap for smaller coding agents Aug 29, 2026
  • Correct SQL answers can still hide broken agent work Aug 27, 2026
  • VBVR-Pro makes visual reasoning a training loop, not a demo Aug 27, 2026
  • Self-improvement audits need a measured null Aug 21, 2026
  • Measure AI precision before you trust AI capability Aug 20, 2026
  • The Self-Improving Agent Demo That Falls Apart When You Shuffle the Tasks Aug 19, 2026
  • CaliBench tests whether video models roll fair dice Aug 18, 2026
  • Stop Making One Prompt Read the Sources and Cast the Vote Aug 17, 2026
  • Auto-research agents need fuzzer-style feedback Aug 11, 2026
  • The TTS Score That Hides Your Voice Model's Real Problems Aug 11, 2026
  • SABRE turns VLM evals into a repeatable stress-test pipeline Aug 10, 2026
  • TutorMoments and the Hard Part of AI Tutoring: Knowing When Not to Answer Aug 8, 2026
  • Agnostic PAC learning gets its optimal bound Aug 7, 2026
  • Benchmarks Need QA Before They Judge Agents Aug 7, 2026
  • Selective Trust: Why RAG Systems Should Learn When to Ignore Their Own Context Aug 7, 2026
  • OpenAI’s cyber eval issue is really a process story Aug 5, 2026
  • What 'Test-Time Scaling' Actually Means When You Read a Benchmark Aug 5, 2026
  • ExtractBench tests document extraction where demos usually fail Aug 3, 2026
  • Computer-use agents need stricter judges, not prettier demos Jul 31, 2026
  • AI research agents can code, but they still can’t judge the work Jul 30, 2026
  • Autonomous research needs a budget scoreboard Jul 28, 2026
  • Safety bounds are becoming probabilities, not vibes Jul 23, 2026
  • When a Model Eval Turns Into an Actual Breach Jul 22, 2026
  • CRAFT Turns Rubrics Into a Fine-Tuning Map Jul 20, 2026
  • Kimi K3 and the narrow win that matters for Next.js builders Jul 18, 2026
  • Safety evals need to test messy language, not just bad behavior Jul 2, 2026
  • Hugging Face model pages get a better eval paper trail Jun 30, 2026
  • Multimodal models still change answers when you shuffle the evidence Jun 25, 2026
  • Self-distillation can make models better on the first try and worse on the fifth Jun 25, 2026
  • Agent Success Rate is the only number that matters when a new model drops May 31, 2026
  • Marketers are still vibe-checking prompts. Frontier devs run evals before lunch. May 30, 2026
  • Stop Vibe-Checking New Models. Build a 50-Prompt Eval Set Instead. May 28, 2026
  • The Frustration Index: A Cheap Eval Most Teams Skip May 20, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public