Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

llm-evaluation

11 posts tagged llm-evaluation.

  • ESPO's fix for prompt bloat: diagnose errors, then stop appending rules Sep 4, 2026
  • The Retrieval-Integration Gap: When Your AI Analyst Reads the Risk and Ignores It Anyway Aug 26, 2026
  • When a Zero-Shot LLM Ties a Random Forest on Travel Behavior Aug 21, 2026
  • Fragile attention paths as a confidence check for grounded QA Aug 12, 2026
  • GeoBenchLLM Tests Whether LLMs Understand Place Aug 10, 2026
  • The Harness Is the Product: What HarnessOpt-Bench Actually Measures Aug 7, 2026
  • The Blind Spot in AI-Text Detectors: Human Writing an LLM Touched Aug 3, 2026
  • When Your RAG Sources Disagree: Kontrast and Cross-Modal Knowledge Auditing Jul 29, 2026
  • The Same Model Name Gave Two Different Answers About Pseudo-Science Jul 27, 2026
  • PoTRE makes the case for heterogeneous test-time reasoning Jul 23, 2026
  • LLMs brainstorm like synthesis machines, not researchers Jul 2, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public