ai-research
45 posts tagged ai-research.
- AI math systems are still optimizing for the wrong proof
- Conversational XAI works best when it is a control surface, not a chatbot
- CRISPR screens need learned experiment pickers, not bigger chatbots
- DataShifts turns distribution shift into an error-budget problem
- MindTopo Tests Whether Vision Models Grasp Topology, Not Just Distances
- Recursive Self-Improvement Has a Roadmap Now, and Most of It Isn't Built
- ConvMem Turns Long-Context Reasoning Into a Tree Instead of a Chain
- Geometry reasoning gets better when the model is not doing every job
- OpenAI training-data accusations need a provenance test
- SG-JEPA points world models at the encoder, not the simulator
- NOAH models patient records as timelines, not snapshots
- Procedural Graphs give agents a memory of what to do, not just what happened
- AlphaGenome Atlas turns genome variants into an AI lookup layer
- OpenAI Says Its Own Researchers Now Lean on Coding Agents. What Does the Data Actually Show?
- Claude, Fermat, and the real value of machine-checkable proofs
- Auxiliary views explain why diverse pre-training data works
- Verbal reinforcement learning is a feedback routing problem
- SWE-Prime argues coding agents need cleaner wins, not more wins
- WikiSkill and the Case for Giving Agents a Memory That Compounds
- Anatomy-informed networks put clinical constraints before more data
- Agent memory works better when skills are small and written in text
- AlphaEvolve nudges the matrix multiplication exponent lower
- AutoSR turns symbolic regression into a research-state search
- PIHF turns prompts into a versioned policy layer
- YOPO makes abstention cheaper by reading the model before it lies
- Alignment baked into pretraining, not bolted on later
- AutoDesign turns paper-to-poster into a harness the agent rewrites itself
- LLMs Know When to Back Off, But Still Guess Too Specifically
- OmniScientist argues that AI scientists need eyes, not just workflows
- CLAUDE.md bloat is a memory problem, not a prompt problem
- Concise answers can weaken reasoning in fused LLM training
- Two prepared policies may be the sweet spot for uncertain MDPs
- AI research agents can code, but they still can’t judge the work
- The Automation Ceiling Nobody Prices In: When Human Participation Is the Product
- No Best Harness: Automated Discovery Systems Fail to Generalize
- Uncertainty metrics should follow the loss, not the other way around
- Vision models are getting the scene right, but not always the gaze
- Leanstral 1.5 points at proofs as a workflow, not a stunt
- DiaLLM separates dialect understanding from dialect writing
- Language critiques are a better training signal than a score, if you can afford them
- LLMs brainstorm like synthesis machines, not researchers
- TraceLab shows coding agents are an infrastructure workload now
- Google's Paper Assistant Wants to Catch Your Math Errors Before a Reviewer Does
- Nash solvers have preferences when the value is identical
- AI math is less about genius than verification