Skip to content
{ ken ashe }
  • Building
  • Topics
  • Blog
  • Building
  • Topics
  • Blog
← Blog / Tags

// tag

agents

62 posts tagged agents.

  • AutoDesign turns paper-to-poster into a harness the agent rewrites itself Aug 14, 2026
  • OmniScientist argues that AI scientists need eyes, not just workflows Aug 14, 2026
  • A Strong Model Can Scaffold a Weak One Without Any Retraining Aug 13, 2026
  • When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL Aug 13, 2026
  • AMIE’s video consult result is about perception, not replacement Aug 11, 2026
  • PsychoAgent makes memory retrieval less purely semantic Aug 10, 2026
  • SkillProx and the Case for Agents That Prune Their Own Playbooks Aug 10, 2026
  • The 2011 Link That Died on Schedule, and What It Says About Agent Memory Aug 9, 2026
  • Benchmarks Need QA Before They Judge Agents Aug 7, 2026
  • Heart-failure feature engineering gets an agent pipeline, not a chatbot Aug 7, 2026
  • RAG For Table-Heavy Reports Needs Search You Can Audit Aug 7, 2026
  • The Harness Is the Product: What HarnessOpt-Bench Actually Measures Aug 7, 2026
  • Video deep research agents need to look before they search Aug 5, 2026
  • LiveMem reframes long-context memory as state continuity Aug 4, 2026
  • DungeonBench puts tactical reasoning where agents usually break Aug 3, 2026
  • ExtractBench tests document extraction where demos usually fail Aug 3, 2026
  • On-policy imitation helps when the student is smaller than the expert Aug 3, 2026
  • Post-training is now the behavior layer Aug 3, 2026
  • The useful question behind a multiplayer agent harness Aug 1, 2026
  • Computer-use agents need stricter judges, not prettier demos Jul 31, 2026
  • When You Count the Tokens, Self-Reflection Loses to Just Sampling More Jul 31, 2026
  • APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups Jul 30, 2026
  • πR² makes robot policies react inside the action chunk Jul 29, 2026
  • Autonomous research needs a budget scoreboard Jul 28, 2026
  • On-Policy Distillation Is Becoming the Default Move for Agent Training Jul 28, 2026
  • Anthropic's SDK Just Named Claude Opus 5 and Added Mid-Stream Tool Swaps Jul 25, 2026
  • GS-Agent Builds 4D Worlds by Driving a Physics Engine, Not Replacing It Jul 24, 2026
  • Training Agents Inside the Harness They Actually Ship With Jul 24, 2026
  • PoTRE makes the case for heterogeneous test-time reasoning Jul 23, 2026
  • O-VAD treats factory video anomalies as object histories Jul 21, 2026
  • Goal prompting is not a solver for NP-hard search Jul 19, 2026
  • ChatGPT’s computer control turns the browser into the new agent runtime Jul 18, 2026
  • Paper Revisions as Training Data: What SciDiagramEdit Gets Right About Figure Editing Jul 17, 2026
  • The One-Shot Trap in Agent Optimization Jul 16, 2026
  • Bonsai 27B makes local agents smaller, not magically smarter Jul 15, 2026
  • Memory Failures Hide Behind Correct Answers: What MemOps Exposes Jul 15, 2026
  • PalmClaw and the Case for Tool-First Mobile Agents Jul 15, 2026
  • Self-repair in small code models may be measuring retry form, not error content Jul 15, 2026
  • What a Model Knows About What It Knows Jul 14, 2026
  • Intelligence Is Still Not the Product Jul 12, 2026
  • What Stampli's 'one person doing four people's work' claim actually shows Jul 11, 2026
  • LangChain’s small July fixes point to bigger agent runtime problems Jul 9, 2026
  • GaP treats robot policies as editable graphs Jul 7, 2026
  • Verification as a Scaling Axis: What the LLM-as-a-Verifier Paper Actually Changes Jul 7, 2026
  • MiniCPM5-1B points at the phone-sized agent layer Jul 5, 2026
  • The Model Got Smarter and My Tool Got Dumber Jul 5, 2026
  • Speaker recognition is a better agent test than another chat demo Jul 3, 2026
  • Anthropic’s Python SDK points to agents as infrastructure, not demos Jul 1, 2026
  • TRIAGE gives agent RL a better target than pass or fail Jul 1, 2026
  • World Models That Edit Their Own Context Instead of Their Weights Jun 30, 2026
  • Agent immunity is the missing layer between alignment and tool use Jun 29, 2026
  • Google's Paper Assistant Wants to Catch Your Math Errors Before a Reviewer Does Jun 29, 2026
  • Nash solvers have preferences when the value is identical Jun 29, 2026
  • Hermes Agent moves the agent demo into the payments layer Jun 28, 2026
  • Hermes turns agent setup into packaging work Jun 28, 2026
  • Promptware Treats Prompt Injection Like an Execution Chain Jun 28, 2026
  • What a one-line LangChain fix tells you about streaming reliability Jun 27, 2026
  • Autodata Turns Synthetic Data Generation Into an Agent You Train Jun 25, 2026
  • Tool-use RL is failing at the brackets, not the tools Jun 25, 2026
  • Codex record-and-replay turns screen demos into reusable skills Jun 23, 2026
  • CUGA's Two Dozen Examples Are the Real Agent Documentation Jun 23, 2026
  • Omio’s AI-native travel bet starts with messy trip planning Jun 23, 2026
  • RSS
  • LinkedIn
  • X
  • GitHub
  • Email
  • Newsroom
  • Media Kit
  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe