ai-research
75 posts tagged ai-research.
- Flow matching gives neural dissimilarity metrics one common frame
- Teaching a Reasoning Model to Know When It's Sure Cuts Its Token Bill
- What LLMs Miss About Haitian Creole, and Why Low-Resource Culture Breaks Evals
- Conspiracy Detection Needs Context, Not Just Better Keywords
- ExplorationBench and the Gap Between Recall and Real Discovery
- MISVO steers frozen language models at inference time
- Authorship verification works better as similarity, not classification
- Clifford-VAE puts pixels into symbolic memory space
- When a Model Gets Math Right but Reads the Same Problem Differently
- A Small Model That Steers a Bigger One's Reasoning
- Long-context models can miss the answer sitting behind nearby noise
- OOD generalization depends on exact mechanisms, not better fits
- Membership Inference Gets an Entropy Correction: Reading the ETD Paper
- OpenAI’s math advisory group is about claim control
- Softmax attention needs an off switch
- Decontamination reports are a promise, not a proof: the case for evaluation-side benchmarks
- A cipher win is not an eval without the working
- ComPO and the Case Against Optimizing the Loss You Wrote Down
- Double descent as an implicit regularization story
- Probabilistic Linear Explanations make interpretability more honest
- What Actually Makes a Tokeniser Good: Search Beats Objective
- When Rewording the Answer Key Reshuffles the Leaderboard
- Fuse tests the weak spot in AI social advice
- LACE compresses speech tokens one codec layer at a time
- ScienceBuddy and the case for training the harness before the model
- Bellman Policy Optimization cuts one moving part from RLVR
- LLM personas fail when opinions have to change
- OptiFlow treats offline RL policy learning as sample matching
- Slip Detection Is Where Robot Hands Stop Dropping Things
- The Skill Router You Already Have: Gavel Reads Routing From a Frozen LLM
- AI math systems are still optimizing for the wrong proof
- Conversational XAI works best when it is a control surface, not a chatbot
- CRISPR screens need learned experiment pickers, not bigger chatbots
- DataShifts turns distribution shift into an error-budget problem
- MindTopo Tests Whether Vision Models Grasp Topology, Not Just Distances
- Recursive Self-Improvement Has a Roadmap Now, and Most of It Isn't Built
- ConvMem Turns Long-Context Reasoning Into a Tree Instead of a Chain
- Geometry reasoning gets better when the model is not doing every job
- OpenAI training-data accusations need a provenance test
- SG-JEPA points world models at the encoder, not the simulator
- NOAH models patient records as timelines, not snapshots
- Procedural Graphs give agents a memory of what to do, not just what happened
- AlphaGenome Atlas turns genome variants into an AI lookup layer
- OpenAI Says Its Own Researchers Now Lean on Coding Agents. What Does the Data Actually Show?
- Claude, Fermat, and the real value of machine-checkable proofs
- Auxiliary views explain why diverse pre-training data works
- Verbal reinforcement learning is a feedback routing problem
- SWE-Prime argues coding agents need cleaner wins, not more wins
- WikiSkill and the Case for Giving Agents a Memory That Compounds
- Anatomy-informed networks put clinical constraints before more data
- Agent memory works better when skills are small and written in text
- AlphaEvolve nudges the matrix multiplication exponent lower
- AutoSR turns symbolic regression into a research-state search
- PIHF turns prompts into a versioned policy layer
- YOPO makes abstention cheaper by reading the model before it lies
- Alignment baked into pretraining, not bolted on later
- AutoDesign turns paper-to-poster into a harness the agent rewrites itself
- LLMs Know When to Back Off, But Still Guess Too Specifically
- OmniScientist argues that AI scientists need eyes, not just workflows
- CLAUDE.md bloat is a memory problem, not a prompt problem
- Concise answers can weaken reasoning in fused LLM training
- Two prepared policies may be the sweet spot for uncertain MDPs
- AI research agents can code, but they still can’t judge the work
- The Automation Ceiling Nobody Prices In: When Human Participation Is the Product
- No Best Harness: Automated Discovery Systems Fail to Generalize
- Uncertainty metrics should follow the loss, not the other way around
- Vision models are getting the scene right, but not always the gaze
- Leanstral 1.5 points at proofs as a workflow, not a stunt
- DiaLLM separates dialect understanding from dialect writing
- Language critiques are a better training signal than a score, if you can afford them
- LLMs brainstorm like synthesis machines, not researchers
- TraceLab shows coding agents are an infrastructure workload now
- Google's Paper Assistant Wants to Catch Your Math Errors Before a Reviewer Does
- Nash solvers have preferences when the value is identical
- AI math is less about genius than verification