Skip to content
{ ken ashe }
  • Building
  • Topics
  • Blog
  • Building
  • Topics
  • Blog
← Blog / Tags

// tag

llm-evaluation

8 posts tagged llm-evaluation.

  • Fragile attention paths as a confidence check for grounded QA Aug 12, 2026
  • GeoBenchLLM Tests Whether LLMs Understand Place Aug 10, 2026
  • The Harness Is the Product: What HarnessOpt-Bench Actually Measures Aug 7, 2026
  • The Blind Spot in AI-Text Detectors: Human Writing an LLM Touched Aug 3, 2026
  • When Your RAG Sources Disagree: Kontrast and Cross-Modal Knowledge Auditing Jul 29, 2026
  • The Same Model Name Gave Two Different Answers About Pseudo-Science Jul 27, 2026
  • PoTRE makes the case for heterogeneous test-time reasoning Jul 23, 2026
  • LLMs brainstorm like synthesis machines, not researchers Jul 2, 2026
  • RSS
  • LinkedIn
  • X
  • GitHub
  • Email
  • Newsroom
  • Media Kit
  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe