model-reliability
80 posts tagged model-reliability.
- Black-box attribute alignment is a sampler, not a fairness wand
- Models may refuse based on who they think you are
- Agent control failures are moving from demo risk to operating risk
- Claude Code’s verification loop is the real coding-agent primitive
- NSA’s reported AI testing spend makes evals look like infrastructure
- Ollaya points at a missing layer for decision models
- When Agents Hack: What the OpenAI-Hugging Face Story Actually Tells Builders
- Audio models do not automatically share speech and text concepts
- Conspiracy Detection Needs Context, Not Just Better Keywords
- MISVO steers frozen language models at inference time
- When a model’s rejection reason changes its next choice
- Australia’s OpenAI agent breach is a disclosure problem, not a sci-fi one
- Authorship verification works better as similarity, not classification
- Multi-agent shutdown sabotage is a real eval target now
- SWE-Flux: The Benchmark That Asks If Coding Models Can Predict What Code Actually Does
- Translation Fine-Tuning Breaks the Controls General Evals Miss
- Fact-checking agents need receipts, not vibes
- Local tool-use evals are measuring your server too
- Long-context models can miss the answer sitting behind nearby noise
- Gemini’s reported company breaches show the AI disclosure gap
- Iterative unalignment tests the failures too rare for normal evals
- Jev shows why science workflows need semantic evals, not just final-answer grading
- Long-horizon agents can learn to cheat the checker
- OOD generalization depends on exact mechanisms, not better fits
- DiaVLo Turns Vision-Language Model Failures Into Named Behaviours
- Multi-hop RAG needs confidence before the answer
- OpenAI wants shared AI standards, but the hard part is enforcement
- OpenAI’s math advisory group is about claim control
- Benchmarks Do Not Settle the Open-Weight Risk Debate
- Decontamination reports are a promise, not a proof: the case for evaluation-side benchmarks
- A cipher win is not an eval without the working
- LangChain 1.4.2 fixes a small but real agent reliability problem
- Marc van der Chijs’s AI banking warning needs plumbing details
- What Two Boring openai-python Patches Reveal About Agent Reliability
- Coding agents overclaim when their work is incomplete
- Jev is a decision model, not another chatbot
- LangChain’s typesafe alpha points at safer model routing
- Microsoft’s AI scraping memo points to a data supply chain problem
- Toxicity Scores Can Miss Sanitized Bias in GPT Outputs
- When a Coding Agent Drives a Robot, Task Success Isn't Safety
- Agent memory helps after the action loop stops breaking
- Double descent as an implicit regularization story
- MUSE tests vision-language models where classroom context gets messy
- Probabilistic Linear Explanations make interpretability more honest
- When Rewording the Answer Key Reshuffles the Leaderboard
- Distillation needs calibration when the teacher is biased
- ENCP gives navigation agents a usable uncertainty check
- Fuse tests the weak spot in AI social advice
- Teaching Models to Say 'I Don't Know' With a Prompt, Not a Retrain
- Zuckerberg’s case for solo AI slowdowns
- AI slowdown promises break without enforcement
- Federated learning gets more practical when privacy and timing are treated together
- K-Bench tests mental health chatbots where generic safety evals do not
- LLM personas fail when opinions have to change
- A vulnerability is not fixed because an AI bot saw it
- Siri as a model router, not a single assistant
- The Thin Claim Behind “No Regulation” for Frontier Model Pacing
- A thin local AI signal still tells builders what to watch
- Real-SWE and the Case for Testing Coding Agents on Code They Have Never Seen
- Search Console Is Not Enough for AI Search Visibility
- When Agents Lie to Pass the Eval
- AI math systems are still optimizing for the wrong proof
- GPT-6 Astra and the problem with moving-target models
- Perplexity Handing GPT-6 Astra Production Access: What OpenAI's Claim Actually Means
- The useful part of a smaller LLM gateway is not the size
- Conversational XAI works best when it is a control surface, not a chatbot
- DataShifts turns distribution shift into an error-budget problem
- RetroThinker lets speech models correct themselves while you are still talking
- The Hallucination Detector That Doesn't Transfer to Your Domain
- CEO clone chatbots expose the limits of personality wrappers
- Geometry reasoning gets better when the model is not doing every job
- Instruction vs. Example: How Vision-Language Models Actually Moderate Content
- Layer-selective unlearning aims at the parts of a model that remember
- OpenAI training-data accusations need a provenance test
- PPC Automation Needs Guardrails, Not Blind Trust
- SG-JEPA points world models at the encoder, not the simulator
- The Benchmark Says GPT-5. Your Users Get a Serving Route.
- OpenAI’s reported Millennium math claim needs proof, not applause
- SPINE Shows Sycophancy Gets Worse When Users Keep Pushing
- What Claude Code's Quiet A/B Test Says About Silent Model Downgrades