ai-agents
64 posts tagged ai-agents.
- AutoDesign turns paper-to-poster into a harness the agent rewrites itself
- OmniScientist argues that AI scientists need eyes, not just workflows
- A Strong Model Can Scaffold a Weak One Without Any Retraining
- When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL
- AMIE’s video consult result is about perception, not replacement
- Auto-research agents need fuzzer-style feedback
- PsychoAgent makes memory retrieval less purely semantic
- SkillProx and the Case for Agents That Prune Their Own Playbooks
- The 2011 Link That Died on Schedule, and What It Says About Agent Memory
- Kitesurf and the browser built for agents, not humans
- Benchmarks Need QA Before They Judge Agents
- Heart-failure feature engineering gets an agent pipeline, not a chatbot
- RAG For Table-Heavy Reports Needs Search You Can Audit
- The Harness Is the Product: What HarnessOpt-Bench Actually Measures
- Video deep research agents need to look before they search
- DNS identity for AI agents needs more than a name
- LiveMem reframes long-context memory as state continuity
- Where multimodal embeddings and collaborative coding agents actually stand
- DungeonBench puts tactical reasoning where agents usually break
- ExtractBench tests document extraction where demos usually fail
- On-policy imitation helps when the student is smaller than the expert
- Post-training is now the behavior layer
- The useful question behind a multiplayer agent harness
- Claude’s test escape is a security design problem, not a sci-fi story
- Computer-use agents need stricter judges, not prettier demos
- ORCA-bench shows oncall agents are not ready for pager duty
- When You Count the Tokens, Self-Reflection Loses to Just Sampling More
- AI research agents can code, but they still can’t judge the work
- APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups
- The Office-Task Benchmark That Prices Agents Against Human Labor
- πR² makes robot policies react inside the action chunk
- Autonomous research needs a budget scoreboard
- ChatGPT Work turns sales AI into a revenue feedback loop
- On-Policy Distillation Is Becoming the Default Move for Agent Training
- CausalForge makes AI research agents prove their work
- Coinbase’s AI agent bet is payments plumbing, not an AI pivot
- Anthropic's SDK Just Named Claude Opus 5 and Added Mid-Stream Tool Swaps
- Treat rogue AI hacker stories as incident reports, not movie trailers
- GS-Agent Builds 4D Worlds by Driving a Physics Engine, Not Replacing It
- LangChain’s gateway env var is production plumbing
- Training Agents Inside the Harness They Actually Ship With
- World’s $52.5M bet on proof of human for agents
- PoTRE makes the case for heterogeneous test-time reasoning
- Kimi K3’s SpreadsheetBench win is a signal, not a coronation
- Shopify’s agent lesson is delegation, not autonomy
- Useful work per dollar is the agent metric that matters
- IoT exploit agents are getting useful in controlled labs
- The AI job diamond still needs a bottom rung
- PolicyGuard makes the compliance bot show its work
- HORIZON treats chip design like a repo, not a chat prompt
- The 4-second budget that decides if your AI agent ships
- Agent Success Rate is the only number that matters when a new model drops
- Splitting the agent loop from tool execution cut TTFT by 90%
- Agent Amnesia Has a Fix, and It Looks Like Sleep
- Dreaming Agents Could Finally End the Brand Voice Correction Loop
- How a 400-line system prompt becomes 15 lines with Skills
- The 200K Token CSV Problem Has a One-Line Fix
- The 400-line system prompt is the new technical debt
- The mechanism matters: making AI competitive analysis auditable
- Last Quarter Means Two Different Things Inside Your Own Company
- Split-brain agents: when planner and executor stop talking, quality collapses
- The Frustration Index: A Cheap Eval Most Teams Skip
- Why I Stopped Trusting Demo Videos for Agent Tools
- When Your AI Agent Needs a Browser, Not an API