building-with-ai
101 posts tagged building-with-ai.
- Black-box attribute alignment is a sampler, not a fairness wand
- Leaves.com shows the new domain decision: sell the name or ship the site
- The useful part of Roetzer’s architect-orchestrator-apprentice AI work model
- When Documentation Doesn't Help Coding Agents: A Negative Result Worth Reading
- Agent control failures are moving from demo risk to operating risk
- Drawgent puts a coding agent on an Excalidraw canvas: what a visual work surface actually changes
- Keeping programming enjoyable when LLMs write the first draft
- Qwen 27B on a 4090 is a useful reminder about local AI demos
- Claude Code’s verification loop is the real coding-agent primitive
- Ollaya points at a missing layer for decision models
- What One Developer Learned From a Month Without AI Tools
- When Agents Hack: What the OpenAI-Hugging Face Story Actually Tells Builders
- AI Agents Make Audience Data Debt More Expensive
- CliffCompaction makes long coding runs cheaper by refusing to summarize
- Local tool-use evals are measuring your server too
- OpenAI’s Python SDK adds model handles before the launch story
- Jev shows why science workflows need semantic evals, not just final-answer grading
- AI adoption is the easy checkbox. Workflow adaptation is the actual job
- AI writing is useful until it starts doing your thinking
- AI posters get better when the model stops being the designer
- Hex, GPT-6 Astra, and the shift from answers to visual reports
- LangChain 1.4.2 fixes a small but real agent reliability problem
- Laya’s Jev claim needs a repo-first reality check
- What Two Boring openai-python Patches Reveal About Agent Reliability
- Jev is a decision model, not another chatbot
- LangChain’s typesafe alpha points at safer model routing
- The Harness Matters as Much as the Model in Coding Agents
- Claude as a transfer-code assistant for domain portfolios
- CareMirror puts caregiver consent at the center of dementia AI
- Slip Detection Is Where Robot Hands Stop Dropping Things
- The Skill Router You Already Have: Gavel Reads Routing From a Frozen LLM
- A vulnerability is not fixed because an AI bot saw it
- JPEG XL Is a Workflow Question, Not a Format War
- Turn repeat marketing work into tools, not just prompts
- What Fyxer's AI inbox assistant gets right about trust
- What llama.cpp v0.4.1 tells you about where local AI is heading
- Real-SWE and the Case for Testing Coding Agents on Code They Have Never Seen
- A Filter That Strips AI Stories From Hacker News, and What It Reveals
- AI sadness is product feedback, not just backlash
- Devin testing its own work with GPT-6 Astra: what's real, what's reported
- Rune going open source is only the start of the diligence
- The RubyGems agent report is a supply-chain warning
- The useful part of a smaller LLM gateway is not the size
- Conversational XAI works best when it is a control surface, not a chatbot
- The marketing AI stack is too full now
- Codex as a lab scout for antimicrobial search
- PPC Automation Needs Guardrails, Not Blind Trust
- Browser agents are becoming practical comment analysts
- OpenAI’s Python SDK gets key expiration controls, not just new image hooks
- Pick the hospital automation work before picking the bot
- Procedural Graphs give agents a memory of what to do, not just what happened
- What ChatGPT Images 2.5 Changes for People Who Actually Ship Images
- AI Made the First Draft Free. The Bottleneck Moved Downstream.
- OpenAI’s journalism program stretches from classrooms to newsrooms
- OpenAI's Python SDK Gets More Honest About Agent Failures
- Agent Harnesses Are Becoming a Token Efficiency Fight
- Local LLMs as the first responder for a compromised PC
- The useful part of calling LLMs a cognitive virus
- The real lesson in a 90% Claude Code token cut
- What LangChain Core 1.6.2 Fixes for Agent Builders
- Repo-distilled skills are the missing middle layer for research agents
- ACToR targets the tokens where repo-level code generation breaks
- AI workflows should create new work, not just faster tasks
- Culture Still Sets the Ceiling on AI Productivity
- No AI Fridays: What a Weekly Ban Reveals About Skill Atrophy
- When LLM memory becomes program analysis
- MCR-Bench shows code review agents still lose the plot
- Anthropic’s Python SDK is tightening the agent plumbing
- Gemini Omni 1.1 Flash needs to prove control, not just speed
- Codex at loveholidays points to the real no-code shift
- OpenAI’s full-stack argument is really an economics argument
- Prime Agent treats the harness as part of the model
- Compliance LLMs need different workflows for passports and DPIAs
- Codex over Claude is a workflow signal, not a verdict
- Munder Difflin and the real work behind AI clone offices
- Claudette and the fight against AI house style
- What LangChain's perplexity 1.4.1 patch says about agent plumbing
- AI literacy is the new fault line: what OpenAI's CodeAI deal actually signals
- Gemini 3.7 Flash and the New Floor for Cheap Models
- AI coding feels like managing a very literal junior engineer
- Alzheimer’s surgery claims need evidence before amplification
- A one-line Hacker News item is not enough context
- What Actually Makes a Claude Code Session Productive
- What OpenAI's v3.1.0 SDK Changelog Tells Us About Its Roadmap
- OlmoEarth embeddings are useful if they survive outside the studio
- RingCentral’s AI-native work story is really about the handoff
- VICBench shows vulnerability detection still needs humans
- AI security review is hitting Bitcoin repos, not just toy code
- Claude’s Bluetooth hint is the right kind of AI assistance
- AI coding cost control is an engineering workflow problem
- Kitesurf and the browser built for agents, not humans
- LLMs as semantic scouts for compiler optimizations
- OpenAI’s education plugins move ChatGPT closer to classroom workflow
- The Financial Advice Chatbots Give You Depends on the Question You Bring
- AI reasoning can look right while taking shortcuts
- Flint points at a missing layer in AI visualization
- Go’s generic collections proposal is really about shared defaults
- The useful question behind a multiplayer agent harness
- RL can teach code models to care about runtime, but the stopwatch is the hard part
- Brolly’s plain-text weather page is an AI product lesson
- Half-Life 2 on HaikuOS and the AI runtime tax