coding-agents
36 posts tagged coding-agents.
- When Documentation Doesn't Help Coding Agents: A Negative Result Worth Reading
- Drawgent puts a coding agent on an Excalidraw canvas: what a visual work surface actually changes
- Keeping programming enjoyable when LLMs write the first draft
- Claude Code’s verification loop is the real coding-agent primitive
- What One Developer Learned From a Month Without AI Tools
- SWE-Flux: The Benchmark That Asks If Coding Models Can Predict What Code Actually Does
- CliffCompaction makes long coding runs cheaper by refusing to summarize
- Coding agents overclaim when their work is incomplete
- The Harness Matters as Much as the Model in Coding Agents
- When a Coding Agent Drives a Robot, Task Success Isn't Safety
- Real-SWE and the Case for Testing Coding Agents on Code They Have Never Seen
- Devin testing its own work with GPT-6 Astra: what's real, what's reported
- Perplexity Handing GPT-6 Astra Production Access: What OpenAI's Claim Actually Means
- Codex as a lab scout for antimicrobial search
- OpenAI Says Its Own Researchers Now Lean on Coding Agents. What Does the Data Actually Show?
- Rust vtables are where AI code reviews get vague
- The real lesson in a 90% Claude Code token cut
- Terminal-Universe turns code-agent logs into reusable sandboxes
- A Minecraft clone is a weak coding-agent test
- GLM, Qwen, and the messy reality of visual coding agents
- Terminal-Bench 4.0 and the eval gap for smaller coding agents
- SWE-Prime argues coding agents need cleaner wins, not more wins
- Coding agents still struggle with whole-repo migrations
- Codex over Claude is a workflow signal, not a verdict
- Local agentic coding at 60 tokens per second is only half the test
- Qwen3.8-27B looks useful as overnight local coding labor
- Qwen’s BASIC ray-tracer demo is really about closed-loop coding
- Oracle’s OpenJDK AI-code ban is really about provenance
- Where multimodal embeddings and collaborative coding agents actually stand
- Instruction following is the local model test benchmarks miss
- MindForge trains coding agents on blank-repo software work
- Coding agents are becoming lab infrastructure
- Claude Fable 5 browser-game demos are really coherence tests
- TestEvo-Bench moves coding-agent evals closer to real maintenance
- TraceLab shows coding agents are an infrastructure workload now
- Coding agents need routers before they need bigger models