Skip to content
{ ken ashe }
  • Building
  • Topics
  • Blog
  • Building
  • Topics
  • Blog
← Blog / Tags

// tag

ai-agents

64 posts tagged ai-agents.

  • AutoDesign turns paper-to-poster into a harness the agent rewrites itself Aug 14, 2026
  • OmniScientist argues that AI scientists need eyes, not just workflows Aug 14, 2026
  • A Strong Model Can Scaffold a Weak One Without Any Retraining Aug 13, 2026
  • When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL Aug 13, 2026
  • AMIE’s video consult result is about perception, not replacement Aug 11, 2026
  • Auto-research agents need fuzzer-style feedback Aug 11, 2026
  • PsychoAgent makes memory retrieval less purely semantic Aug 10, 2026
  • SkillProx and the Case for Agents That Prune Their Own Playbooks Aug 10, 2026
  • The 2011 Link That Died on Schedule, and What It Says About Agent Memory Aug 9, 2026
  • Kitesurf and the browser built for agents, not humans Aug 8, 2026
  • Benchmarks Need QA Before They Judge Agents Aug 7, 2026
  • Heart-failure feature engineering gets an agent pipeline, not a chatbot Aug 7, 2026
  • RAG For Table-Heavy Reports Needs Search You Can Audit Aug 7, 2026
  • The Harness Is the Product: What HarnessOpt-Bench Actually Measures Aug 7, 2026
  • Video deep research agents need to look before they search Aug 5, 2026
  • DNS identity for AI agents needs more than a name Aug 4, 2026
  • LiveMem reframes long-context memory as state continuity Aug 4, 2026
  • Where multimodal embeddings and collaborative coding agents actually stand Aug 4, 2026
  • DungeonBench puts tactical reasoning where agents usually break Aug 3, 2026
  • ExtractBench tests document extraction where demos usually fail Aug 3, 2026
  • On-policy imitation helps when the student is smaller than the expert Aug 3, 2026
  • Post-training is now the behavior layer Aug 3, 2026
  • The useful question behind a multiplayer agent harness Aug 1, 2026
  • Claude’s test escape is a security design problem, not a sci-fi story Jul 31, 2026
  • Computer-use agents need stricter judges, not prettier demos Jul 31, 2026
  • ORCA-bench shows oncall agents are not ready for pager duty Jul 31, 2026
  • When You Count the Tokens, Self-Reflection Loses to Just Sampling More Jul 31, 2026
  • AI research agents can code, but they still can’t judge the work Jul 30, 2026
  • APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups Jul 30, 2026
  • The Office-Task Benchmark That Prices Agents Against Human Labor Jul 30, 2026
  • πR² makes robot policies react inside the action chunk Jul 29, 2026
  • Autonomous research needs a budget scoreboard Jul 28, 2026
  • ChatGPT Work turns sales AI into a revenue feedback loop Jul 28, 2026
  • On-Policy Distillation Is Becoming the Default Move for Agent Training Jul 28, 2026
  • CausalForge makes AI research agents prove their work Jul 27, 2026
  • Coinbase’s AI agent bet is payments plumbing, not an AI pivot Jul 27, 2026
  • Anthropic's SDK Just Named Claude Opus 5 and Added Mid-Stream Tool Swaps Jul 25, 2026
  • Treat rogue AI hacker stories as incident reports, not movie trailers Jul 25, 2026
  • GS-Agent Builds 4D Worlds by Driving a Physics Engine, Not Replacing It Jul 24, 2026
  • LangChain’s gateway env var is production plumbing Jul 24, 2026
  • Training Agents Inside the Harness They Actually Ship With Jul 24, 2026
  • World’s $52.5M bet on proof of human for agents Jul 24, 2026
  • PoTRE makes the case for heterogeneous test-time reasoning Jul 23, 2026
  • Kimi K3’s SpreadsheetBench win is a signal, not a coronation Jul 19, 2026
  • Shopify’s agent lesson is delegation, not autonomy Jul 18, 2026
  • Useful work per dollar is the agent metric that matters Jul 15, 2026
  • IoT exploit agents are getting useful in controlled labs Jul 13, 2026
  • The AI job diamond still needs a bottom rung Jul 5, 2026
  • PolicyGuard makes the compliance bot show its work Jul 1, 2026
  • HORIZON treats chip design like a repo, not a chat prompt Jun 29, 2026
  • The 4-second budget that decides if your AI agent ships Jun 3, 2026
  • Agent Success Rate is the only number that matters when a new model drops May 31, 2026
  • Splitting the agent loop from tool execution cut TTFT by 90% May 27, 2026
  • Agent Amnesia Has a Fix, and It Looks Like Sleep May 26, 2026
  • Dreaming Agents Could Finally End the Brand Voice Correction Loop May 26, 2026
  • How a 400-line system prompt becomes 15 lines with Skills May 26, 2026
  • The 200K Token CSV Problem Has a One-Line Fix May 24, 2026
  • The 400-line system prompt is the new technical debt May 23, 2026
  • The mechanism matters: making AI competitive analysis auditable May 22, 2026
  • Last Quarter Means Two Different Things Inside Your Own Company May 21, 2026
  • Split-brain agents: when planner and executor stop talking, quality collapses May 21, 2026
  • The Frustration Index: A Cheap Eval Most Teams Skip May 20, 2026
  • Why I Stopped Trusting Demo Videos for Agent Tools May 20, 2026
  • When Your AI Agent Needs a Browser, Not an API May 15, 2026
  • RSS
  • LinkedIn
  • X
  • GitHub
  • Email
  • Newsroom
  • Media Kit
  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe