evals
76 posts tagged evals.
- Flow matching gives neural dissimilarity metrics one common frame
- What LLMs Miss About Haitian Creole, and Why Low-Resource Culture Breaks Evals
- NSA’s reported AI testing spend makes evals look like infrastructure
- ExplorationBench and the Gap Between Recall and Real Discovery
- When a model’s rejection reason changes its next choice
- SWE-Flux: The Benchmark That Asks If Coding Models Can Predict What Code Actually Does
- Translation Fine-Tuning Breaks the Controls General Evals Miss
- When a Model Gets Math Right but Reads the Same Problem Differently
- Long-context models can miss the answer sitting behind nearby noise
- DolphinBench Tests Agent Memory by Task, Not Trivia
- Iterative unalignment tests the failures too rare for normal evals
- Jev shows why science workflows need semantic evals, not just final-answer grading
- DiaVLo Turns Vision-Language Model Failures Into Named Behaviours
- Multi-hop RAG needs confidence before the answer
- OpenAI wants shared AI standards, but the hard part is enforcement
- OpenAI’s math advisory group is about claim control
- Decontamination reports are a promise, not a proof: the case for evaluation-side benchmarks
- A cipher win is not an eval without the working
- Toxicity Scores Can Miss Sanitized Bias in GPT Outputs
- Probabilistic Linear Explanations make interpretability more honest
- When Rewording the Answer Key Reshuffles the Leaderboard
- Teaching Models to Say 'I Don't Know' With a Prompt, Not a Retrain
- K-Bench tests mental health chatbots where generic safety evals do not
- LLM personas fail when opinions have to change
- When Agents Lie to Pass the Eval
- AI math systems are still optimizing for the wrong proof
- GPT-6 Astra and the problem with moving-target models
- DataShifts turns distribution shift into an error-budget problem
- GEO needs exposure data before it can claim revenue
- The Hallucination Detector That Doesn't Transfer to Your Domain
- Instruction vs. Example: How Vision-Language Models Actually Moderate Content
- OpenAI training-data accusations need a provenance test
- The Benchmark Says GPT-5. Your Users Get a Serving Route.
- SPINE Shows Sycophancy Gets Worse When Users Keep Pushing
- LexFlip Shows Legal Meaning Metrics Are Still Token Counters
- Hosted LLM judges are not fixed instruments
- AICOME works best when the missing variable is narrow
- Bare assertions are a weak spot for medical reasoning models
- When a Language Model's Odds Don't Add Up
- Grading Coding Agents on How They Solve, Not Just Whether They Pass
- Reward choice changes how LLM forecasters are wrong
- Synthetic data needs a weight limit
- A Minecraft clone is a weak coding-agent test
- Terminal-Bench 4.0 and the eval gap for smaller coding agents
- Correct SQL answers can still hide broken agent work
- VBVR-Pro makes visual reasoning a training loop, not a demo
- Self-improvement audits need a measured null
- Measure AI precision before you trust AI capability
- The Self-Improving Agent Demo That Falls Apart When You Shuffle the Tasks
- CaliBench tests whether video models roll fair dice
- Stop Making One Prompt Read the Sources and Cast the Vote
- Auto-research agents need fuzzer-style feedback
- The TTS Score That Hides Your Voice Model's Real Problems
- SABRE turns VLM evals into a repeatable stress-test pipeline
- TutorMoments and the Hard Part of AI Tutoring: Knowing When Not to Answer
- Agnostic PAC learning gets its optimal bound
- Benchmarks Need QA Before They Judge Agents
- Selective Trust: Why RAG Systems Should Learn When to Ignore Their Own Context
- OpenAI’s cyber eval issue is really a process story
- What 'Test-Time Scaling' Actually Means When You Read a Benchmark
- ExtractBench tests document extraction where demos usually fail
- Computer-use agents need stricter judges, not prettier demos
- AI research agents can code, but they still can’t judge the work
- Autonomous research needs a budget scoreboard
- Safety bounds are becoming probabilities, not vibes
- When a Model Eval Turns Into an Actual Breach
- CRAFT Turns Rubrics Into a Fine-Tuning Map
- Kimi K3 and the narrow win that matters for Next.js builders
- Safety evals need to test messy language, not just bad behavior
- Hugging Face model pages get a better eval paper trail
- Multimodal models still change answers when you shuffle the evidence
- Self-distillation can make models better on the first try and worse on the fifth
- Agent Success Rate is the only number that matters when a new model drops
- Marketers are still vibe-checking prompts. Frontier devs run evals before lunch.
- Stop Vibe-Checking New Models. Build a 50-Prompt Eval Set Instead.
- The Frustration Index: A Cheap Eval Most Teams Skip