agents
93 posts tagged agents.
- OpenAI's Python SDK Gets More Honest About Agent Failures
- How a 9B model learned to chain Korean government APIs by actually calling them
- GPT-6 Astra's Big Claim: Agentic Skill Meets Alignment
- What LangChain Core 1.6.2 Fixes for Agent Builders
- Compiling a Prompt Into a Small Neural Function You Can Version
- ESPO's fix for prompt bloat: diagnose errors, then stop appending rules
- Repo-distilled skills are the missing middle layer for research agents
- Proactive writing agents need timing more than autocomplete
- SAGE Uses a Big Model as a Coach, Not a Crutch
- Verbal reinforcement learning is a feedback routing problem
- Gemini’s agentic video framing shifts the hard part from seeing to checking
- Small dialogue agents need repair loops, not just bigger training
- What Happens When You Let AI Agents Build Their Own Society
- WikiSkill and the Case for Giving Agents a Memory That Compounds
- Anthropic’s Python SDK is tightening the agent plumbing
- Recuris and the Case for Memory That Rewrites Itself
- Prime Agent treats the harness as part of the model
- Verification becomes the operator loop for AI-built silicon
- A 36-node DGX Spark homelab points at agent infrastructure, not just bigger inference
- Munder Difflin and the real work behind AI clone offices
- Treat the Texas AI hacking story as a disclosure systems test
- What LangChain's perplexity 1.4.1 patch says about agent plumbing
- Agent memory works better when skills are small and written in text
- Task Model Induction turns messy screen traces into reusable agent skills
- When a Zero-Shot LLM Ties a Random Forest on Travel Behavior
- SPADE Makes the Training Environment a Thing the Model Learns to Build
- The Self-Improving Agent Demo That Falls Apart When You Shuffle the Tasks
- AutoSR turns symbolic regression into a research-state search
- PIHF turns prompts into a versioned policy layer
- AI coding feels like managing a very literal junior engineer
- What Actually Makes a Claude Code Session Productive
- AutoDesign turns paper-to-poster into a harness the agent rewrites itself
- OmniScientist argues that AI scientists need eyes, not just workflows
- A Strong Model Can Scaffold a Weak One Without Any Retraining
- When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL
- AMIE’s video consult result is about perception, not replacement
- PsychoAgent makes memory retrieval less purely semantic
- SkillProx and the Case for Agents That Prune Their Own Playbooks
- The 2011 Link That Died on Schedule, and What It Says About Agent Memory
- Benchmarks Need QA Before They Judge Agents
- Heart-failure feature engineering gets an agent pipeline, not a chatbot
- RAG For Table-Heavy Reports Needs Search You Can Audit
- The Harness Is the Product: What HarnessOpt-Bench Actually Measures
- Video deep research agents need to look before they search
- LiveMem reframes long-context memory as state continuity
- DungeonBench puts tactical reasoning where agents usually break
- ExtractBench tests document extraction where demos usually fail
- On-policy imitation helps when the student is smaller than the expert
- Post-training is now the behavior layer
- The useful question behind a multiplayer agent harness
- Computer-use agents need stricter judges, not prettier demos
- When You Count the Tokens, Self-Reflection Loses to Just Sampling More
- APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups
- πR² makes robot policies react inside the action chunk
- Autonomous research needs a budget scoreboard
- On-Policy Distillation Is Becoming the Default Move for Agent Training
- Anthropic's SDK Just Named Claude Opus 5 and Added Mid-Stream Tool Swaps
- GS-Agent Builds 4D Worlds by Driving a Physics Engine, Not Replacing It
- Training Agents Inside the Harness They Actually Ship With
- PoTRE makes the case for heterogeneous test-time reasoning
- O-VAD treats factory video anomalies as object histories
- Goal prompting is not a solver for NP-hard search
- ChatGPT’s computer control turns the browser into the new agent runtime
- Paper Revisions as Training Data: What SciDiagramEdit Gets Right About Figure Editing
- The One-Shot Trap in Agent Optimization
- Bonsai 27B makes local agents smaller, not magically smarter
- Memory Failures Hide Behind Correct Answers: What MemOps Exposes
- PalmClaw and the Case for Tool-First Mobile Agents
- Self-repair in small code models may be measuring retry form, not error content
- What a Model Knows About What It Knows
- Intelligence Is Still Not the Product
- What Stampli's 'one person doing four people's work' claim actually shows
- LangChain’s small July fixes point to bigger agent runtime problems
- GaP treats robot policies as editable graphs
- Verification as a Scaling Axis: What the LLM-as-a-Verifier Paper Actually Changes
- MiniCPM5-1B points at the phone-sized agent layer
- The Model Got Smarter and My Tool Got Dumber
- Speaker recognition is a better agent test than another chat demo
- Anthropic’s Python SDK points to agents as infrastructure, not demos
- TRIAGE gives agent RL a better target than pass or fail
- World Models That Edit Their Own Context Instead of Their Weights
- Agent immunity is the missing layer between alignment and tool use
- Google's Paper Assistant Wants to Catch Your Math Errors Before a Reviewer Does
- Nash solvers have preferences when the value is identical
- Hermes Agent moves the agent demo into the payments layer
- Hermes turns agent setup into packaging work
- Promptware Treats Prompt Injection Like an Execution Chain
- What a one-line LangChain fix tells you about streaming reliability
- Autodata Turns Synthetic Data Generation Into an Agent You Train
- Tool-use RL is failing at the brackets, not the tools
- Codex record-and-replay turns screen demos into reusable skills
- CUGA's Two Dozen Examples Are the Real Agent Documentation
- Omio’s AI-native travel bet starts with messy trip planning