Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

model-reliability

80 posts tagged model-reliability.

  • Black-box attribute alignment is a sampler, not a fairness wand Sep 28, 2026
  • Models may refuse based on who they think you are Sep 28, 2026
  • Agent control failures are moving from demo risk to operating risk Sep 27, 2026
  • Claude Code’s verification loop is the real coding-agent primitive Sep 26, 2026
  • NSA’s reported AI testing spend makes evals look like infrastructure Sep 26, 2026
  • Ollaya points at a missing layer for decision models Sep 26, 2026
  • When Agents Hack: What the OpenAI-Hugging Face Story Actually Tells Builders Sep 26, 2026
  • Audio models do not automatically share speech and text concepts Sep 25, 2026
  • Conspiracy Detection Needs Context, Not Just Better Keywords Sep 25, 2026
  • MISVO steers frozen language models at inference time Sep 25, 2026
  • When a model’s rejection reason changes its next choice Sep 25, 2026
  • Australia’s OpenAI agent breach is a disclosure problem, not a sci-fi one Sep 24, 2026
  • Authorship verification works better as similarity, not classification Sep 24, 2026
  • Multi-agent shutdown sabotage is a real eval target now Sep 24, 2026
  • SWE-Flux: The Benchmark That Asks If Coding Models Can Predict What Code Actually Does Sep 24, 2026
  • Translation Fine-Tuning Breaks the Controls General Evals Miss Sep 24, 2026
  • Fact-checking agents need receipts, not vibes Sep 23, 2026
  • Local tool-use evals are measuring your server too Sep 23, 2026
  • Long-context models can miss the answer sitting behind nearby noise Sep 23, 2026
  • Gemini’s reported company breaches show the AI disclosure gap Sep 22, 2026
  • Iterative unalignment tests the failures too rare for normal evals Sep 22, 2026
  • Jev shows why science workflows need semantic evals, not just final-answer grading Sep 22, 2026
  • Long-horizon agents can learn to cheat the checker Sep 22, 2026
  • OOD generalization depends on exact mechanisms, not better fits Sep 22, 2026
  • DiaVLo Turns Vision-Language Model Failures Into Named Behaviours Sep 21, 2026
  • Multi-hop RAG needs confidence before the answer Sep 21, 2026
  • OpenAI wants shared AI standards, but the hard part is enforcement Sep 21, 2026
  • OpenAI’s math advisory group is about claim control Sep 21, 2026
  • Benchmarks Do Not Settle the Open-Weight Risk Debate Sep 20, 2026
  • Decontamination reports are a promise, not a proof: the case for evaluation-side benchmarks Sep 20, 2026
  • A cipher win is not an eval without the working Sep 19, 2026
  • LangChain 1.4.2 fixes a small but real agent reliability problem Sep 19, 2026
  • Marc van der Chijs’s AI banking warning needs plumbing details Sep 19, 2026
  • What Two Boring openai-python Patches Reveal About Agent Reliability Sep 19, 2026
  • Coding agents overclaim when their work is incomplete Sep 18, 2026
  • Jev is a decision model, not another chatbot Sep 18, 2026
  • LangChain’s typesafe alpha points at safer model routing Sep 18, 2026
  • Microsoft’s AI scraping memo points to a data supply chain problem Sep 18, 2026
  • Toxicity Scores Can Miss Sanitized Bias in GPT Outputs Sep 18, 2026
  • When a Coding Agent Drives a Robot, Task Success Isn't Safety Sep 18, 2026
  • Agent memory helps after the action loop stops breaking Sep 17, 2026
  • Double descent as an implicit regularization story Sep 17, 2026
  • MUSE tests vision-language models where classroom context gets messy Sep 17, 2026
  • Probabilistic Linear Explanations make interpretability more honest Sep 17, 2026
  • When Rewording the Answer Key Reshuffles the Leaderboard Sep 17, 2026
  • Distillation needs calibration when the teacher is biased Sep 16, 2026
  • ENCP gives navigation agents a usable uncertainty check Sep 16, 2026
  • Fuse tests the weak spot in AI social advice Sep 16, 2026
  • Teaching Models to Say 'I Don't Know' With a Prompt, Not a Retrain Sep 16, 2026
  • Zuckerberg’s case for solo AI slowdowns Sep 16, 2026
  • AI slowdown promises break without enforcement Sep 15, 2026
  • Federated learning gets more practical when privacy and timing are treated together Sep 15, 2026
  • K-Bench tests mental health chatbots where generic safety evals do not Sep 15, 2026
  • LLM personas fail when opinions have to change Sep 15, 2026
  • A vulnerability is not fixed because an AI bot saw it Sep 14, 2026
  • Siri as a model router, not a single assistant Sep 14, 2026
  • The Thin Claim Behind “No Regulation” for Frontier Model Pacing Sep 14, 2026
  • A thin local AI signal still tells builders what to watch Sep 13, 2026
  • Real-SWE and the Case for Testing Coding Agents on Code They Have Never Seen Sep 13, 2026
  • Search Console Is Not Enough for AI Search Visibility Sep 13, 2026
  • When Agents Lie to Pass the Eval Sep 13, 2026
  • AI math systems are still optimizing for the wrong proof Sep 12, 2026
  • GPT-6 Astra and the problem with moving-target models Sep 12, 2026
  • Perplexity Handing GPT-6 Astra Production Access: What OpenAI's Claim Actually Means Sep 12, 2026
  • The useful part of a smaller LLM gateway is not the size Sep 12, 2026
  • Conversational XAI works best when it is a control surface, not a chatbot Sep 11, 2026
  • DataShifts turns distribution shift into an error-budget problem Sep 11, 2026
  • RetroThinker lets speech models correct themselves while you are still talking Sep 11, 2026
  • The Hallucination Detector That Doesn't Transfer to Your Domain Sep 11, 2026
  • CEO clone chatbots expose the limits of personality wrappers Sep 10, 2026
  • Geometry reasoning gets better when the model is not doing every job Sep 10, 2026
  • Instruction vs. Example: How Vision-Language Models Actually Moderate Content Sep 10, 2026
  • Layer-selective unlearning aims at the parts of a model that remember Sep 10, 2026
  • OpenAI training-data accusations need a provenance test Sep 10, 2026
  • PPC Automation Needs Guardrails, Not Blind Trust Sep 10, 2026
  • SG-JEPA points world models at the encoder, not the simulator Sep 10, 2026
  • The Benchmark Says GPT-5. Your Users Get a Serving Route. Sep 10, 2026
  • OpenAI’s reported Millennium math claim needs proof, not applause Sep 9, 2026
  • SPINE Shows Sycophancy Gets Worse When Users Keep Pushing Sep 9, 2026
  • What Claude Code's Quiet A/B Test Says About Silent Model Downgrades Aug 23, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public