llm-evaluation
11 posts tagged llm-evaluation.
- ESPO's fix for prompt bloat: diagnose errors, then stop appending rules
- The Retrieval-Integration Gap: When Your AI Analyst Reads the Risk and Ignores It Anyway
- When a Zero-Shot LLM Ties a Random Forest on Travel Behavior
- Fragile attention paths as a confidence check for grounded QA
- GeoBenchLLM Tests Whether LLMs Understand Place
- The Harness Is the Product: What HarnessOpt-Bench Actually Measures
- The Blind Spot in AI-Text Detectors: Human Writing an LLM Touched
- When Your RAG Sources Disagree: Kontrast and Cross-Modal Knowledge Auditing
- The Same Model Name Gave Two Different Answers About Pseudo-Science
- PoTRE makes the case for heterogeneous test-time reasoning
- LLMs brainstorm like synthesis machines, not researchers