evals
25 posts tagged evals.
- Auto-research agents need fuzzer-style feedback
- The TTS Score That Hides Your Voice Model's Real Problems
- SABRE turns VLM evals into a repeatable stress-test pipeline
- TutorMoments and the Hard Part of AI Tutoring: Knowing When Not to Answer
- Agnostic PAC learning gets its optimal bound
- Benchmarks Need QA Before They Judge Agents
- Selective Trust: Why RAG Systems Should Learn When to Ignore Their Own Context
- OpenAI’s cyber eval issue is really a process story
- What 'Test-Time Scaling' Actually Means When You Read a Benchmark
- ExtractBench tests document extraction where demos usually fail
- Computer-use agents need stricter judges, not prettier demos
- AI research agents can code, but they still can’t judge the work
- Autonomous research needs a budget scoreboard
- Safety bounds are becoming probabilities, not vibes
- When a Model Eval Turns Into an Actual Breach
- CRAFT Turns Rubrics Into a Fine-Tuning Map
- Kimi K3 and the narrow win that matters for Next.js builders
- Safety evals need to test messy language, not just bad behavior
- Hugging Face model pages get a better eval paper trail
- Multimodal models still change answers when you shuffle the evidence
- Self-distillation can make models better on the first try and worse on the fifth
- Agent Success Rate is the only number that matters when a new model drops
- Marketers are still vibe-checking prompts. Frontier devs run evals before lunch.
- Stop Vibe-Checking New Models. Build a 50-Prompt Eval Set Instead.
- The Frustration Index: A Cheap Eval Most Teams Skip