Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

ai-agents

121 posts tagged ai-agents.

  • Devin testing its own work with GPT-6 Astra: what's real, what's reported Sep 12, 2026
  • Perplexity Handing GPT-6 Astra Production Access: What OpenAI's Claim Actually Means Sep 12, 2026
  • The RubyGems agent report is a supply-chain warning Sep 12, 2026
  • An 86% Bitcoin quantum benchmark cut is not an 86% attack Sep 11, 2026
  • Recursive Self-Improvement Has a Roadmap Now, and Most of It Isn't Built Sep 11, 2026
  • CEO clone chatbots expose the limits of personality wrappers Sep 10, 2026
  • Browser agents are becoming practical comment analysts Sep 9, 2026
  • Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet Sep 9, 2026
  • OpenAI’s reported Millennium math claim needs proof, not applause Sep 9, 2026
  • Procedural Graphs give agents a memory of what to do, not just what happened Sep 9, 2026
  • OpenAI's Python SDK Gets More Honest About Agent Failures Sep 8, 2026
  • How a 9B model learned to chain Korean government APIs by actually calling them Sep 7, 2026
  • When You Upgrade the Model, Does the Agent's Memory Come With It? Sep 7, 2026
  • Agent Harnesses Are Becoming a Token Efficiency Fight Sep 6, 2026
  • AI can help with PCB design, but it cannot own signoff Sep 5, 2026
  • GPT-6 Astra's Big Claim: Agentic Skill Meets Alignment Sep 5, 2026
  • What LangChain Core 1.6.2 Fixes for Agent Builders Sep 5, 2026
  • Compiling a Prompt Into a Small Neural Function You Can Version Sep 4, 2026
  • ESPO's fix for prompt bloat: diagnose errors, then stop appending rules Sep 4, 2026
  • The OpenAI agent hack report is a disclosure problem first Sep 4, 2026
  • Repo-distilled skills are the missing middle layer for research agents Sep 3, 2026
  • Web agents need useful predictions, not prettier page snapshots Sep 3, 2026
  • AI agent interviews are a workflow fix, not a content strategy Sep 2, 2026
  • Proactive writing agents need timing more than autocomplete Sep 2, 2026
  • SAGE Uses a Big Model as a Coach, Not a Crutch Sep 2, 2026
  • Verbal reinforcement learning is a feedback routing problem Sep 2, 2026
  • Gemini’s agentic video framing shifts the hard part from seeing to checking Sep 1, 2026
  • The Evaluate-First Turn in AI Research Agents Sep 1, 2026
  • Small dialogue agents need repair loops, not just bigger training Aug 31, 2026
  • What Happens When You Let AI Agents Build Their Own Society Aug 30, 2026
  • Agent payments are a product problem before they are a currency problem Aug 29, 2026
  • Treat the smart TV like an untrusted computer Aug 29, 2026
  • WikiSkill and the Case for Giving Agents a Memory That Compounds Aug 28, 2026
  • Anthropic’s Python SDK is tightening the agent plumbing Aug 27, 2026
  • Correct SQL answers can still hide broken agent work Aug 27, 2026
  • SwarmWorld makes a case for agents that coordinate through artifacts Aug 27, 2026
  • Claude Code as a domain renewal research assistant Aug 26, 2026
  • Recuris and the Case for Memory That Rewrites Itself Aug 26, 2026
  • The Retrieval-Integration Gap: When Your AI Analyst Reads the Risk and Ignores It Anyway Aug 26, 2026
  • Prime Agent treats the harness as part of the model Aug 25, 2026
  • Gemini and Apex point prediction markets toward brokerage plumbing Aug 24, 2026
  • Verification becomes the operator loop for AI-built silicon Aug 24, 2026
  • A 36-node DGX Spark homelab points at agent infrastructure, not just bigger inference Aug 23, 2026
  • Munder Difflin and the real work behind AI clone offices Aug 23, 2026
  • Treat the Texas AI hacking story as a disclosure systems test Aug 23, 2026
  • What LangChain's perplexity 1.4.1 patch says about agent plumbing Aug 22, 2026
  • Agent memory works better when skills are small and written in text Aug 21, 2026
  • Binance Agent OS puts trading agents behind user-controlled gates Aug 21, 2026
  • Task Model Induction turns messy screen traces into reusable agent skills Aug 21, 2026
  • What AI4AI-Bench Says About Recursive Self-Improvement Right Now Aug 21, 2026
  • When a Zero-Shot LLM Ties a Random Forest on Travel Behavior Aug 21, 2026
  • SPADE Makes the Training Environment a Thing the Model Learns to Build Aug 20, 2026
  • The Self-Improving Agent Demo That Falls Apart When You Shuffle the Tasks Aug 19, 2026
  • AutoSR turns symbolic regression into a research-state search Aug 18, 2026
  • PIHF turns prompts into a versioned policy layer Aug 18, 2026
  • AI coding feels like managing a very literal junior engineer Aug 16, 2026
  • What Actually Makes a Claude Code Session Productive Aug 15, 2026
  • AutoDesign turns paper-to-poster into a harness the agent rewrites itself Aug 14, 2026
  • OmniScientist argues that AI scientists need eyes, not just workflows Aug 14, 2026
  • A Strong Model Can Scaffold a Weak One Without Any Retraining Aug 13, 2026
  • When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL Aug 13, 2026
  • AMIE’s video consult result is about perception, not replacement Aug 11, 2026
  • Auto-research agents need fuzzer-style feedback Aug 11, 2026
  • PsychoAgent makes memory retrieval less purely semantic Aug 10, 2026
  • SkillProx and the Case for Agents That Prune Their Own Playbooks Aug 10, 2026
  • The 2011 Link That Died on Schedule, and What It Says About Agent Memory Aug 9, 2026
  • Kitesurf and the browser built for agents, not humans Aug 8, 2026
  • Benchmarks Need QA Before They Judge Agents Aug 7, 2026
  • Heart-failure feature engineering gets an agent pipeline, not a chatbot Aug 7, 2026
  • RAG For Table-Heavy Reports Needs Search You Can Audit Aug 7, 2026
  • The Harness Is the Product: What HarnessOpt-Bench Actually Measures Aug 7, 2026
  • Video deep research agents need to look before they search Aug 5, 2026
  • DNS identity for AI agents needs more than a name Aug 4, 2026
  • LiveMem reframes long-context memory as state continuity Aug 4, 2026
  • Where multimodal embeddings and collaborative coding agents actually stand Aug 4, 2026
  • DungeonBench puts tactical reasoning where agents usually break Aug 3, 2026
  • ExtractBench tests document extraction where demos usually fail Aug 3, 2026
  • On-policy imitation helps when the student is smaller than the expert Aug 3, 2026
  • Post-training is now the behavior layer Aug 3, 2026
  • The useful question behind a multiplayer agent harness Aug 1, 2026
  • Claude’s test escape is a security design problem, not a sci-fi story Jul 31, 2026
  • Computer-use agents need stricter judges, not prettier demos Jul 31, 2026
  • ORCA-bench shows oncall agents are not ready for pager duty Jul 31, 2026
  • When You Count the Tokens, Self-Reflection Loses to Just Sampling More Jul 31, 2026
  • AI research agents can code, but they still can’t judge the work Jul 30, 2026
  • APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups Jul 30, 2026
  • The Office-Task Benchmark That Prices Agents Against Human Labor Jul 30, 2026
  • πR² makes robot policies react inside the action chunk Jul 29, 2026
  • Autonomous research needs a budget scoreboard Jul 28, 2026
  • ChatGPT Work turns sales AI into a revenue feedback loop Jul 28, 2026
  • On-Policy Distillation Is Becoming the Default Move for Agent Training Jul 28, 2026
  • CausalForge makes AI research agents prove their work Jul 27, 2026
  • Coinbase’s AI agent bet is payments plumbing, not an AI pivot Jul 27, 2026
  • Anthropic's SDK Just Named Claude Opus 5 and Added Mid-Stream Tool Swaps Jul 25, 2026
  • Treat rogue AI hacker stories as incident reports, not movie trailers Jul 25, 2026
  • GS-Agent Builds 4D Worlds by Driving a Physics Engine, Not Replacing It Jul 24, 2026
  • LangChain’s gateway env var is production plumbing Jul 24, 2026
  • Training Agents Inside the Harness They Actually Ship With Jul 24, 2026
  • World’s $52.5M bet on proof of human for agents Jul 24, 2026
  • PoTRE makes the case for heterogeneous test-time reasoning Jul 23, 2026
  • Kimi K3’s SpreadsheetBench win is a signal, not a coronation Jul 19, 2026
  • Shopify’s agent lesson is delegation, not autonomy Jul 18, 2026
  • Useful work per dollar is the agent metric that matters Jul 15, 2026
  • IoT exploit agents are getting useful in controlled labs Jul 13, 2026
  • The AI job diamond still needs a bottom rung Jul 5, 2026
  • PolicyGuard makes the compliance bot show its work Jul 1, 2026
  • HORIZON treats chip design like a repo, not a chat prompt Jun 29, 2026
  • The 4-second budget that decides if your AI agent ships Jun 3, 2026
  • Agent Success Rate is the only number that matters when a new model drops May 31, 2026
  • Splitting the agent loop from tool execution cut TTFT by 90% May 27, 2026
  • Agent Amnesia Has a Fix, and It Looks Like Sleep May 26, 2026
  • Dreaming Agents Could Finally End the Brand Voice Correction Loop May 26, 2026
  • How a 400-line system prompt becomes 15 lines with Skills May 26, 2026
  • The 200K Token CSV Problem Has a One-Line Fix May 24, 2026
  • The 400-line system prompt is the new technical debt May 23, 2026
  • The mechanism matters: making AI competitive analysis auditable May 22, 2026
  • Last Quarter Means Two Different Things Inside Your Own Company May 21, 2026
  • Split-brain agents: when planner and executor stop talking, quality collapses May 21, 2026
  • The Frustration Index: A Cheap Eval Most Teams Skip May 20, 2026
  • Why I Stopped Trusting Demo Videos for Agent Tools May 20, 2026
  • When Your AI Agent Needs a Browser, Not an API May 15, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public