Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

agents

93 posts tagged agents.

  • OpenAI's Python SDK Gets More Honest About Agent Failures Sep 8, 2026
  • How a 9B model learned to chain Korean government APIs by actually calling them Sep 7, 2026
  • GPT-6 Astra's Big Claim: Agentic Skill Meets Alignment Sep 5, 2026
  • What LangChain Core 1.6.2 Fixes for Agent Builders Sep 5, 2026
  • Compiling a Prompt Into a Small Neural Function You Can Version Sep 4, 2026
  • ESPO's fix for prompt bloat: diagnose errors, then stop appending rules Sep 4, 2026
  • Repo-distilled skills are the missing middle layer for research agents Sep 3, 2026
  • Proactive writing agents need timing more than autocomplete Sep 2, 2026
  • SAGE Uses a Big Model as a Coach, Not a Crutch Sep 2, 2026
  • Verbal reinforcement learning is a feedback routing problem Sep 2, 2026
  • Gemini’s agentic video framing shifts the hard part from seeing to checking Sep 1, 2026
  • Small dialogue agents need repair loops, not just bigger training Aug 31, 2026
  • What Happens When You Let AI Agents Build Their Own Society Aug 30, 2026
  • WikiSkill and the Case for Giving Agents a Memory That Compounds Aug 28, 2026
  • Anthropic’s Python SDK is tightening the agent plumbing Aug 27, 2026
  • Recuris and the Case for Memory That Rewrites Itself Aug 26, 2026
  • Prime Agent treats the harness as part of the model Aug 25, 2026
  • Verification becomes the operator loop for AI-built silicon Aug 24, 2026
  • A 36-node DGX Spark homelab points at agent infrastructure, not just bigger inference Aug 23, 2026
  • Munder Difflin and the real work behind AI clone offices Aug 23, 2026
  • Treat the Texas AI hacking story as a disclosure systems test Aug 23, 2026
  • What LangChain's perplexity 1.4.1 patch says about agent plumbing Aug 22, 2026
  • Agent memory works better when skills are small and written in text Aug 21, 2026
  • Task Model Induction turns messy screen traces into reusable agent skills Aug 21, 2026
  • When a Zero-Shot LLM Ties a Random Forest on Travel Behavior Aug 21, 2026
  • SPADE Makes the Training Environment a Thing the Model Learns to Build Aug 20, 2026
  • The Self-Improving Agent Demo That Falls Apart When You Shuffle the Tasks Aug 19, 2026
  • AutoSR turns symbolic regression into a research-state search Aug 18, 2026
  • PIHF turns prompts into a versioned policy layer Aug 18, 2026
  • AI coding feels like managing a very literal junior engineer Aug 16, 2026
  • What Actually Makes a Claude Code Session Productive Aug 15, 2026
  • AutoDesign turns paper-to-poster into a harness the agent rewrites itself Aug 14, 2026
  • OmniScientist argues that AI scientists need eyes, not just workflows Aug 14, 2026
  • A Strong Model Can Scaffold a Weak One Without Any Retraining Aug 13, 2026
  • When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL Aug 13, 2026
  • AMIE’s video consult result is about perception, not replacement Aug 11, 2026
  • PsychoAgent makes memory retrieval less purely semantic Aug 10, 2026
  • SkillProx and the Case for Agents That Prune Their Own Playbooks Aug 10, 2026
  • The 2011 Link That Died on Schedule, and What It Says About Agent Memory Aug 9, 2026
  • Benchmarks Need QA Before They Judge Agents Aug 7, 2026
  • Heart-failure feature engineering gets an agent pipeline, not a chatbot Aug 7, 2026
  • RAG For Table-Heavy Reports Needs Search You Can Audit Aug 7, 2026
  • The Harness Is the Product: What HarnessOpt-Bench Actually Measures Aug 7, 2026
  • Video deep research agents need to look before they search Aug 5, 2026
  • LiveMem reframes long-context memory as state continuity Aug 4, 2026
  • DungeonBench puts tactical reasoning where agents usually break Aug 3, 2026
  • ExtractBench tests document extraction where demos usually fail Aug 3, 2026
  • On-policy imitation helps when the student is smaller than the expert Aug 3, 2026
  • Post-training is now the behavior layer Aug 3, 2026
  • The useful question behind a multiplayer agent harness Aug 1, 2026
  • Computer-use agents need stricter judges, not prettier demos Jul 31, 2026
  • When You Count the Tokens, Self-Reflection Loses to Just Sampling More Jul 31, 2026
  • APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups Jul 30, 2026
  • πR² makes robot policies react inside the action chunk Jul 29, 2026
  • Autonomous research needs a budget scoreboard Jul 28, 2026
  • On-Policy Distillation Is Becoming the Default Move for Agent Training Jul 28, 2026
  • Anthropic's SDK Just Named Claude Opus 5 and Added Mid-Stream Tool Swaps Jul 25, 2026
  • GS-Agent Builds 4D Worlds by Driving a Physics Engine, Not Replacing It Jul 24, 2026
  • Training Agents Inside the Harness They Actually Ship With Jul 24, 2026
  • PoTRE makes the case for heterogeneous test-time reasoning Jul 23, 2026
  • O-VAD treats factory video anomalies as object histories Jul 21, 2026
  • Goal prompting is not a solver for NP-hard search Jul 19, 2026
  • ChatGPT’s computer control turns the browser into the new agent runtime Jul 18, 2026
  • Paper Revisions as Training Data: What SciDiagramEdit Gets Right About Figure Editing Jul 17, 2026
  • The One-Shot Trap in Agent Optimization Jul 16, 2026
  • Bonsai 27B makes local agents smaller, not magically smarter Jul 15, 2026
  • Memory Failures Hide Behind Correct Answers: What MemOps Exposes Jul 15, 2026
  • PalmClaw and the Case for Tool-First Mobile Agents Jul 15, 2026
  • Self-repair in small code models may be measuring retry form, not error content Jul 15, 2026
  • What a Model Knows About What It Knows Jul 14, 2026
  • Intelligence Is Still Not the Product Jul 12, 2026
  • What Stampli's 'one person doing four people's work' claim actually shows Jul 11, 2026
  • LangChain’s small July fixes point to bigger agent runtime problems Jul 9, 2026
  • GaP treats robot policies as editable graphs Jul 7, 2026
  • Verification as a Scaling Axis: What the LLM-as-a-Verifier Paper Actually Changes Jul 7, 2026
  • MiniCPM5-1B points at the phone-sized agent layer Jul 5, 2026
  • The Model Got Smarter and My Tool Got Dumber Jul 5, 2026
  • Speaker recognition is a better agent test than another chat demo Jul 3, 2026
  • Anthropic’s Python SDK points to agents as infrastructure, not demos Jul 1, 2026
  • TRIAGE gives agent RL a better target than pass or fail Jul 1, 2026
  • World Models That Edit Their Own Context Instead of Their Weights Jun 30, 2026
  • Agent immunity is the missing layer between alignment and tool use Jun 29, 2026
  • Google's Paper Assistant Wants to Catch Your Math Errors Before a Reviewer Does Jun 29, 2026
  • Nash solvers have preferences when the value is identical Jun 29, 2026
  • Hermes Agent moves the agent demo into the payments layer Jun 28, 2026
  • Hermes turns agent setup into packaging work Jun 28, 2026
  • Promptware Treats Prompt Injection Like an Execution Chain Jun 28, 2026
  • What a one-line LangChain fix tells you about streaming reliability Jun 27, 2026
  • Autodata Turns Synthetic Data Generation Into an Agent You Train Jun 25, 2026
  • Tool-use RL is failing at the brackets, not the tools Jun 25, 2026
  • Codex record-and-replay turns screen demos into reusable skills Jun 23, 2026
  • CUGA's Two Dozen Examples Are the Real Agent Documentation Jun 23, 2026
  • Omio’s AI-native travel bet starts with messy trip planning Jun 23, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public