agents
62 posts tagged agents.
- AutoDesign turns paper-to-poster into a harness the agent rewrites itself
- OmniScientist argues that AI scientists need eyes, not just workflows
- A Strong Model Can Scaffold a Weak One Without Any Retraining
- When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL
- AMIE’s video consult result is about perception, not replacement
- PsychoAgent makes memory retrieval less purely semantic
- SkillProx and the Case for Agents That Prune Their Own Playbooks
- The 2011 Link That Died on Schedule, and What It Says About Agent Memory
- Benchmarks Need QA Before They Judge Agents
- Heart-failure feature engineering gets an agent pipeline, not a chatbot
- RAG For Table-Heavy Reports Needs Search You Can Audit
- The Harness Is the Product: What HarnessOpt-Bench Actually Measures
- Video deep research agents need to look before they search
- LiveMem reframes long-context memory as state continuity
- DungeonBench puts tactical reasoning where agents usually break
- ExtractBench tests document extraction where demos usually fail
- On-policy imitation helps when the student is smaller than the expert
- Post-training is now the behavior layer
- The useful question behind a multiplayer agent harness
- Computer-use agents need stricter judges, not prettier demos
- When You Count the Tokens, Self-Reflection Loses to Just Sampling More
- APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups
- πR² makes robot policies react inside the action chunk
- Autonomous research needs a budget scoreboard
- On-Policy Distillation Is Becoming the Default Move for Agent Training
- Anthropic's SDK Just Named Claude Opus 5 and Added Mid-Stream Tool Swaps
- GS-Agent Builds 4D Worlds by Driving a Physics Engine, Not Replacing It
- Training Agents Inside the Harness They Actually Ship With
- PoTRE makes the case for heterogeneous test-time reasoning
- O-VAD treats factory video anomalies as object histories
- Goal prompting is not a solver for NP-hard search
- ChatGPT’s computer control turns the browser into the new agent runtime
- Paper Revisions as Training Data: What SciDiagramEdit Gets Right About Figure Editing
- The One-Shot Trap in Agent Optimization
- Bonsai 27B makes local agents smaller, not magically smarter
- Memory Failures Hide Behind Correct Answers: What MemOps Exposes
- PalmClaw and the Case for Tool-First Mobile Agents
- Self-repair in small code models may be measuring retry form, not error content
- What a Model Knows About What It Knows
- Intelligence Is Still Not the Product
- What Stampli's 'one person doing four people's work' claim actually shows
- LangChain’s small July fixes point to bigger agent runtime problems
- GaP treats robot policies as editable graphs
- Verification as a Scaling Axis: What the LLM-as-a-Verifier Paper Actually Changes
- MiniCPM5-1B points at the phone-sized agent layer
- The Model Got Smarter and My Tool Got Dumber
- Speaker recognition is a better agent test than another chat demo
- Anthropic’s Python SDK points to agents as infrastructure, not demos
- TRIAGE gives agent RL a better target than pass or fail
- World Models That Edit Their Own Context Instead of Their Weights
- Agent immunity is the missing layer between alignment and tool use
- Google's Paper Assistant Wants to Catch Your Math Errors Before a Reviewer Does
- Nash solvers have preferences when the value is identical
- Hermes Agent moves the agent demo into the payments layer
- Hermes turns agent setup into packaging work
- Promptware Treats Prompt Injection Like an Execution Chain
- What a one-line LangChain fix tells you about streaming reliability
- Autodata Turns Synthetic Data Generation Into an Agent You Train
- Tool-use RL is failing at the brackets, not the tools
- Codex record-and-replay turns screen demos into reusable skills
- CUGA's Two Dozen Examples Are the Real Agent Documentation
- Omio’s AI-native travel bet starts with messy trip planning