ai-agents
121 posts tagged ai-agents.
- Devin testing its own work with GPT-6 Astra: what's real, what's reported
- Perplexity Handing GPT-6 Astra Production Access: What OpenAI's Claim Actually Means
- The RubyGems agent report is a supply-chain warning
- An 86% Bitcoin quantum benchmark cut is not an 86% attack
- Recursive Self-Improvement Has a Roadmap Now, and Most of It Isn't Built
- CEO clone chatbots expose the limits of personality wrappers
- Browser agents are becoming practical comment analysts
- Can AI Agents Do Their Own Interpretability Research? A New Benchmark Says Not Yet
- OpenAI’s reported Millennium math claim needs proof, not applause
- Procedural Graphs give agents a memory of what to do, not just what happened
- OpenAI's Python SDK Gets More Honest About Agent Failures
- How a 9B model learned to chain Korean government APIs by actually calling them
- When You Upgrade the Model, Does the Agent's Memory Come With It?
- Agent Harnesses Are Becoming a Token Efficiency Fight
- AI can help with PCB design, but it cannot own signoff
- GPT-6 Astra's Big Claim: Agentic Skill Meets Alignment
- What LangChain Core 1.6.2 Fixes for Agent Builders
- Compiling a Prompt Into a Small Neural Function You Can Version
- ESPO's fix for prompt bloat: diagnose errors, then stop appending rules
- The OpenAI agent hack report is a disclosure problem first
- Repo-distilled skills are the missing middle layer for research agents
- Web agents need useful predictions, not prettier page snapshots
- AI agent interviews are a workflow fix, not a content strategy
- Proactive writing agents need timing more than autocomplete
- SAGE Uses a Big Model as a Coach, Not a Crutch
- Verbal reinforcement learning is a feedback routing problem
- Gemini’s agentic video framing shifts the hard part from seeing to checking
- The Evaluate-First Turn in AI Research Agents
- Small dialogue agents need repair loops, not just bigger training
- What Happens When You Let AI Agents Build Their Own Society
- Agent payments are a product problem before they are a currency problem
- Treat the smart TV like an untrusted computer
- WikiSkill and the Case for Giving Agents a Memory That Compounds
- Anthropic’s Python SDK is tightening the agent plumbing
- Correct SQL answers can still hide broken agent work
- SwarmWorld makes a case for agents that coordinate through artifacts
- Claude Code as a domain renewal research assistant
- Recuris and the Case for Memory That Rewrites Itself
- The Retrieval-Integration Gap: When Your AI Analyst Reads the Risk and Ignores It Anyway
- Prime Agent treats the harness as part of the model
- Gemini and Apex point prediction markets toward brokerage plumbing
- Verification becomes the operator loop for AI-built silicon
- A 36-node DGX Spark homelab points at agent infrastructure, not just bigger inference
- Munder Difflin and the real work behind AI clone offices
- Treat the Texas AI hacking story as a disclosure systems test
- What LangChain's perplexity 1.4.1 patch says about agent plumbing
- Agent memory works better when skills are small and written in text
- Binance Agent OS puts trading agents behind user-controlled gates
- Task Model Induction turns messy screen traces into reusable agent skills
- What AI4AI-Bench Says About Recursive Self-Improvement Right Now
- When a Zero-Shot LLM Ties a Random Forest on Travel Behavior
- SPADE Makes the Training Environment a Thing the Model Learns to Build
- The Self-Improving Agent Demo That Falls Apart When You Shuffle the Tasks
- AutoSR turns symbolic regression into a research-state search
- PIHF turns prompts into a versioned policy layer
- AI coding feels like managing a very literal junior engineer
- What Actually Makes a Claude Code Session Productive
- AutoDesign turns paper-to-poster into a harness the agent rewrites itself
- OmniScientist argues that AI scientists need eyes, not just workflows
- A Strong Model Can Scaffold a Weak One Without Any Retraining
- When Your User Simulator Is Secretly One Person: Simulator Collapse in Multi-Agent RL
- AMIE’s video consult result is about perception, not replacement
- Auto-research agents need fuzzer-style feedback
- PsychoAgent makes memory retrieval less purely semantic
- SkillProx and the Case for Agents That Prune Their Own Playbooks
- The 2011 Link That Died on Schedule, and What It Says About Agent Memory
- Kitesurf and the browser built for agents, not humans
- Benchmarks Need QA Before They Judge Agents
- Heart-failure feature engineering gets an agent pipeline, not a chatbot
- RAG For Table-Heavy Reports Needs Search You Can Audit
- The Harness Is the Product: What HarnessOpt-Bench Actually Measures
- Video deep research agents need to look before they search
- DNS identity for AI agents needs more than a name
- LiveMem reframes long-context memory as state continuity
- Where multimodal embeddings and collaborative coding agents actually stand
- DungeonBench puts tactical reasoning where agents usually break
- ExtractBench tests document extraction where demos usually fail
- On-policy imitation helps when the student is smaller than the expert
- Post-training is now the behavior layer
- The useful question behind a multiplayer agent harness
- Claude’s test escape is a security design problem, not a sci-fi story
- Computer-use agents need stricter judges, not prettier demos
- ORCA-bench shows oncall agents are not ready for pager duty
- When You Count the Tokens, Self-Reflection Loses to Just Sampling More
- AI research agents can code, but they still can’t judge the work
- APEX-Accounting: The 2.6% Number That Should Scare AI Bookkeeping Startups
- The Office-Task Benchmark That Prices Agents Against Human Labor
- πR² makes robot policies react inside the action chunk
- Autonomous research needs a budget scoreboard
- ChatGPT Work turns sales AI into a revenue feedback loop
- On-Policy Distillation Is Becoming the Default Move for Agent Training
- CausalForge makes AI research agents prove their work
- Coinbase’s AI agent bet is payments plumbing, not an AI pivot
- Anthropic's SDK Just Named Claude Opus 5 and Added Mid-Stream Tool Swaps
- Treat rogue AI hacker stories as incident reports, not movie trailers
- GS-Agent Builds 4D Worlds by Driving a Physics Engine, Not Replacing It
- LangChain’s gateway env var is production plumbing
- Training Agents Inside the Harness They Actually Ship With
- World’s $52.5M bet on proof of human for agents
- PoTRE makes the case for heterogeneous test-time reasoning
- Kimi K3’s SpreadsheetBench win is a signal, not a coronation
- Shopify’s agent lesson is delegation, not autonomy
- Useful work per dollar is the agent metric that matters
- IoT exploit agents are getting useful in controlled labs
- The AI job diamond still needs a bottom rung
- PolicyGuard makes the compliance bot show its work
- HORIZON treats chip design like a repo, not a chat prompt
- The 4-second budget that decides if your AI agent ships
- Agent Success Rate is the only number that matters when a new model drops
- Splitting the agent loop from tool execution cut TTFT by 90%
- Agent Amnesia Has a Fix, and It Looks Like Sleep
- Dreaming Agents Could Finally End the Brand Voice Correction Loop
- How a 400-line system prompt becomes 15 lines with Skills
- The 200K Token CSV Problem Has a One-Line Fix
- The 400-line system prompt is the new technical debt
- The mechanism matters: making AI competitive analysis auditable
- Last Quarter Means Two Different Things Inside Your Own Company
- Split-brain agents: when planner and executor stop talking, quality collapses
- The Frustration Index: A Cheap Eval Most Teams Skip
- Why I Stopped Trusting Demo Videos for Agent Tools
- When Your AI Agent Needs a Browser, Not an API