Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

ai-agents

162 posts tagged ai-agents.

  • Australia’s AI agent hearing is really about incident disclosure Sep 28, 2026
  • The useful part of Roetzer’s architect-orchestrator-apprentice AI work model Sep 28, 2026
  • When Documentation Doesn't Help Coding Agents: A Negative Result Worth Reading Sep 28, 2026
  • Agent control failures are moving from demo risk to operating risk Sep 27, 2026
  • Drawgent puts a coding agent on an Excalidraw canvas: what a visual work surface actually changes Sep 27, 2026
  • AI agents cut the estimated cost of quantum-safe Bitcoin transactions Sep 26, 2026
  • Ollaya points at a missing layer for decision models Sep 26, 2026
  • When Agents Hack: What the OpenAI-Hugging Face Story Actually Tells Builders Sep 26, 2026
  • AI Agents Make Audience Data Debt More Expensive Sep 25, 2026
  • Bitcoin Lightning joins x402, and agent payments get a little more real Sep 25, 2026
  • Conspiracy Detection Needs Context, Not Just Better Keywords Sep 25, 2026
  • ExplorationBench and the Gap Between Recall and Real Discovery Sep 25, 2026
  • Australia’s OpenAI agent breach is a disclosure problem, not a sci-fi one Sep 24, 2026
  • Multi-agent shutdown sabotage is a real eval target now Sep 24, 2026
  • The World Model That Edits the Agent's Mind, Not the Environment Sep 24, 2026
  • AI agents need payment rails before they need crypto hype Sep 23, 2026
  • Fact-checking agents need receipts, not vibes Sep 23, 2026
  • Known puts domain identity into the agent stack Sep 23, 2026
  • Local tool-use evals are measuring your server too Sep 23, 2026
  • Speaker-Centered Memory: Why Group Chats Break Your AI Agent Sep 23, 2026
  • Critical-State RL trains the one agent call that actually matters Sep 22, 2026
  • DolphinBench Tests Agent Memory by Task, Not Trivia Sep 22, 2026
  • Gemini’s reported company breaches show the AI disclosure gap Sep 22, 2026
  • Long-horizon agents can learn to cheat the checker Sep 22, 2026
  • Where Harness Self-Improvement Actually Stands in Late 2026 Sep 22, 2026
  • Full-Duplex Voice With Tool Calls: What NemotronLabs VoiceChat Actually Ships Sep 21, 2026
  • Hex, GPT-6 Astra, and the shift from answers to visual reports Sep 19, 2026
  • LangChain 1.4.2 fixes a small but real agent reliability problem Sep 19, 2026
  • What Two Boring openai-python Patches Reveal About Agent Reliability Sep 19, 2026
  • LangChain’s typesafe alpha points at safer model routing Sep 18, 2026
  • The Harness Matters as Much as the Model in Coding Agents Sep 18, 2026
  • When a Coding Agent Drives a Robot, Task Success Isn't Safety Sep 18, 2026
  • When the Teacher Knows to Quit: RetireOPD and Distillation for Agents Sep 18, 2026
  • Agent memory helps after the action loop stops breaking Sep 17, 2026
  • ENCP gives navigation agents a usable uncertainty check Sep 16, 2026
  • ScienceBuddy and the case for training the harness before the model Sep 16, 2026
  • The Skill Router You Already Have: Gavel Reads Routing From a Frozen LLM Sep 15, 2026
  • A vulnerability is not fixed because an AI bot saw it Sep 14, 2026
  • Siri as a model router, not a single assistant Sep 14, 2026
  • What Fyxer's AI inbox assistant gets right about trust Sep 14, 2026
  • When Agents Lie to Pass the Eval Sep 13, 2026
  • Devin testing its own work with GPT-6 Astra: what's real, what's reported Sep 12, 2026
  • Perplexity Handing GPT-6 Astra Production Access: What OpenAI's Claim Actually Means Sep 12, 2026
  • The RubyGems agent report is a supply-chain warning Sep 12, 2026
  • An 86% Bitcoin quantum benchmark cut is not an 86% attack Sep 11, 2026
  • Recursive Self-Improvement Has a Roadmap Now, and Most of It Isn't Built Sep 11, 2026
  • CEO clone chatbots expose the limits of personality wrappers Sep 10, 2026
  • Browser agents are becoming practical comment analysts Sep 9, 2026
  • Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet Sep 9, 2026
  • OpenAI’s reported Millennium math claim needs proof, not applause Sep 9, 2026
  • Procedural Graphs give agents a memory of what to do, not just what happened Sep 9, 2026
  • OpenAI's Python SDK Gets More Honest About Agent Failures Sep 8, 2026
  • How a 9B model learned to chain Korean government APIs by actually calling them Sep 7, 2026
  • When You Upgrade the Model, Does the Agent's Memory Come With It? Sep 7, 2026
  • Agent Harnesses Are Becoming a Token Efficiency Fight Sep 6, 2026
  • AI can help with PCB design, but it cannot own signoff Sep 5, 2026
  • GPT-6 Astra's Big Claim: Agentic Skill Meets Alignment Sep 5, 2026
  • What LangChain Core 1.6.2 Fixes for Agent Builders Sep 5, 2026
  • Compiling a Prompt Into a Small Neural Function You Can Version Sep 4, 2026
  • ESPO's fix for prompt bloat: diagnose errors, then stop appending rules Sep 4, 2026
  • The OpenAI agent hack report is a disclosure problem first Sep 4, 2026
  • Repo-distilled skills are the missing middle layer for research agents Sep 3, 2026
  • Web agents need useful predictions, not prettier page snapshots Sep 3, 2026
  • AI agent interviews are a workflow fix, not a content strategy Sep 2, 2026
  • Proactive writing agents need timing more than autocomplete Sep 2, 2026
  • SAGE Uses a Big Model as a Coach, Not a Crutch Sep 2, 2026
  • Verbal reinforcement learning is a feedback routing problem Sep 2, 2026
  • Gemini’s agentic video framing shifts the hard part from seeing to checking Sep 1, 2026
  • The Evaluate-First Turn in AI Research Agents Sep 1, 2026
  • Small dialogue agents need repair loops, not just bigger training Aug 31, 2026
  • What Happens When You Let AI Agents Build Their Own Society Aug 30, 2026
  • Agent payments are a product problem before they are a currency problem Aug 29, 2026
  • Treat the smart TV like an untrusted computer Aug 29, 2026
  • WikiSkill and the Case for Giving Agents a Memory That Compounds Aug 28, 2026
  • Anthropic’s Python SDK is tightening the agent plumbing Aug 27, 2026
  • Correct SQL answers can still hide broken agent work Aug 27, 2026
  • SwarmWorld makes a case for agents that coordinate through artifacts Aug 27, 2026
  • Claude Code as a domain renewal research assistant Aug 26, 2026
  • Recuris and the Case for Memory That Rewrites Itself Aug 26, 2026
  • The Retrieval-Integration Gap: When Your AI Analyst Reads the Risk and Ignores It Anyway Aug 26, 2026
  • Prime Agent treats the harness as part of the model Aug 25, 2026
  • Gemini and Apex point prediction markets toward brokerage plumbing Aug 24, 2026
  • Verification becomes the operator loop for AI-built silicon Aug 24, 2026
  • A 36-node DGX Spark homelab points at agent infrastructure, not just bigger inference Aug 23, 2026
  • Munder Difflin and the real work behind AI clone offices Aug 23, 2026
  • Treat the Texas AI hacking story as a disclosure systems test Aug 23, 2026
  • What LangChain's perplexity 1.4.1 patch says about agent plumbing Aug 22, 2026
  • Agent memory works better when skills are small and written in text Aug 21, 2026
  • Binance Agent OS puts trading agents behind user-controlled gates Aug 21, 2026
  • Task Model Induction turns messy screen traces into reusable agent skills Aug 21, 2026
  • What AI4AI-Bench Says About Recursive Self-Improvement Right Now Aug 21, 2026
  • When a Zero-Shot LLM Ties a Random Forest on Travel Behavior Aug 21, 2026
  • SPADE Makes the Training Environment a Thing the Model Learns to Build Aug 20, 2026
  • The Self-Improving Agent Demo That Falls Apart When You Shuffle the Tasks Aug 19, 2026
  • AutoSR turns symbolic regression into a research-state search Aug 18, 2026
  • PIHF turns prompts into a versioned policy layer Aug 18, 2026
  • AI coding feels like managing a very literal junior engineer Aug 16, 2026
  • What Actually Makes a Claude Code Session Productive Aug 15, 2026
  • AutoDesign turns paper-to-poster into a harness the agent rewrites itself Aug 14, 2026
  • OmniScientist argues that AI scientists need eyes, not just workflows Aug 14, 2026
  • A Strong Model Can Scaffold a Weak One Without Any Retraining Aug 13, 2026
  • When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL Aug 13, 2026
  • AMIE’s video consult result is about perception, not replacement Aug 11, 2026
  • Auto-research agents need fuzzer-style feedback Aug 11, 2026
  • PsychoAgent makes memory retrieval less purely semantic Aug 10, 2026
  • SkillProx and the Case for Agents That Prune Their Own Playbooks Aug 10, 2026
  • The 2011 Link That Died on Schedule, and What It Says About Agent Memory Aug 9, 2026
  • Kitesurf and the browser built for agents, not humans Aug 8, 2026
  • Benchmarks Need QA Before They Judge Agents Aug 7, 2026
  • Heart-failure feature engineering gets an agent pipeline, not a chatbot Aug 7, 2026
  • RAG For Table-Heavy Reports Needs Search You Can Audit Aug 7, 2026
  • The Harness Is the Product: What HarnessOpt-Bench Actually Measures Aug 7, 2026
  • Video deep research agents need to look before they search Aug 5, 2026
  • DNS identity for AI agents needs more than a name Aug 4, 2026
  • LiveMem reframes long-context memory as state continuity Aug 4, 2026
  • Where multimodal embeddings and collaborative coding agents actually stand Aug 4, 2026
  • DungeonBench puts tactical reasoning where agents usually break Aug 3, 2026
  • ExtractBench tests document extraction where demos usually fail Aug 3, 2026
  • On-policy imitation helps when the student is smaller than the expert Aug 3, 2026
  • Post-training is now the behavior layer Aug 3, 2026
  • The useful question behind a multiplayer agent harness Aug 1, 2026
  • Claude’s test escape is a security design problem, not a sci-fi story Jul 31, 2026
  • Computer-use agents need stricter judges, not prettier demos Jul 31, 2026
  • ORCA-bench shows oncall agents are not ready for pager duty Jul 31, 2026
  • When You Count the Tokens, Self-Reflection Loses to Just Sampling More Jul 31, 2026
  • AI research agents can code, but they still can’t judge the work Jul 30, 2026
  • APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups Jul 30, 2026
  • The Office-Task Benchmark That Prices Agents Against Human Labor Jul 30, 2026
  • πR² makes robot policies react inside the action chunk Jul 29, 2026
  • Autonomous research needs a budget scoreboard Jul 28, 2026
  • ChatGPT Work turns sales AI into a revenue feedback loop Jul 28, 2026
  • On-Policy Distillation Is Becoming the Default Move for Agent Training Jul 28, 2026
  • CausalForge makes AI research agents prove their work Jul 27, 2026
  • Coinbase’s AI agent bet is payments plumbing, not an AI pivot Jul 27, 2026
  • Anthropic's SDK Just Named Claude Opus 5 and Added Mid-Stream Tool Swaps Jul 25, 2026
  • Treat rogue AI hacker stories as incident reports, not movie trailers Jul 25, 2026
  • GS-Agent Builds 4D Worlds by Driving a Physics Engine, Not Replacing It Jul 24, 2026
  • LangChain’s gateway env var is production plumbing Jul 24, 2026
  • Training Agents Inside the Harness They Actually Ship With Jul 24, 2026
  • World’s $52.5M bet on proof of human for agents Jul 24, 2026
  • PoTRE makes the case for heterogeneous test-time reasoning Jul 23, 2026
  • Kimi K3’s SpreadsheetBench win is a signal, not a coronation Jul 19, 2026
  • Shopify’s agent lesson is delegation, not autonomy Jul 18, 2026
  • Useful work per dollar is the agent metric that matters Jul 15, 2026
  • IoT exploit agents are getting useful in controlled labs Jul 13, 2026
  • The AI job diamond still needs a bottom rung Jul 5, 2026
  • PolicyGuard makes the compliance bot show its work Jul 1, 2026
  • HORIZON treats chip design like a repo, not a chat prompt Jun 29, 2026
  • The 4-second budget that decides if your AI agent ships Jun 3, 2026
  • Agent Success Rate is the only number that matters when a new model drops May 31, 2026
  • Splitting the agent loop from tool execution cut TTFT by 90% May 27, 2026
  • Agent Amnesia Has a Fix, and It Looks Like Sleep May 26, 2026
  • Dreaming Agents Could Finally End the Brand Voice Correction Loop May 26, 2026
  • How a 400-line system prompt becomes 15 lines with Skills May 26, 2026
  • The 200K Token CSV Problem Has a One-Line Fix May 24, 2026
  • The 400-line system prompt is the new technical debt May 23, 2026
  • The mechanism matters: making AI competitive analysis auditable May 22, 2026
  • Last Quarter Means Two Different Things Inside Your Own Company May 21, 2026
  • Split-brain agents: when planner and executor stop talking, quality collapses May 21, 2026
  • The Frustration Index: A Cheap Eval Most Teams Skip May 20, 2026
  • Why I Stopped Trusting Demo Videos for Agent Tools May 20, 2026
  • When Your AI Agent Needs a Browser, Not an API May 15, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public