llm-evaluation
8 posts tagged llm-evaluation.
- Fragile attention paths as a confidence check for grounded QA
- GeoBenchLLM Tests Whether LLMs Understand Place
- The Harness Is the Product: What HarnessOpt-Bench Actually Measures
- The Blind Spot in AI-Text Detectors: Human Writing an LLM Touched
- When Your RAG Sources Disagree: Kontrast and Cross-Modal Knowledge Auditing
- The Same Model Name Gave Two Different Answers About Pseudo-Science
- PoTRE makes the case for heterogeneous test-time reasoning
- LLMs brainstorm like synthesis machines, not researchers