Skip to content
{ ken ashe }
  • Building
  • Topics
  • Blog
  • Building
  • Topics
  • Blog
← Blog / Tags

// tag

llm-evaluation

8 posts tagged llm-evaluation.

  • Fragile attention paths as a confidence check for grounded QA Aug 12, 2026
  • GeoBenchLLM Tests Whether LLMs Understand Place Aug 10, 2026
  • The Harness Is the Product: What HarnessOpt-Bench Actually Measures Aug 7, 2026
  • The Blind Spot in AI-Text Detectors: Human Writing an LLM Touched Aug 3, 2026
  • When Your RAG Sources Disagree: Kontrast and Cross-Modal Knowledge Auditing Jul 29, 2026
  • The Same Model Name Gave Two Different Answers About Pseudo-Science Jul 27, 2026
  • PoTRE makes the case for heterogeneous test-time reasoning Jul 23, 2026
  • LLMs brainstorm like synthesis machines, not researchers Jul 2, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Site

  • Building
  • Topics
  • Blog
  • Newsroom
  • Media Kit
  • Lucky Domains

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public