model-training
38 posts tagged model-training.
- Free local AI models are not charity
- Clifford-VAE puts pixels into symbolic memory space
- UECR-GRPO treats the teacher as evidence, not an oracle
- Critical-State RL trains the one agent call that actually matters
- onPanda turns alignment feedback into token-level steering
- Where Harness Self-Improvement Actually Stands in Late 2026
- Membership Inference Gets an Entropy Correction: Reading the ETD Paper
- Mini-AGI makes local training the interesting part
- Softmax attention needs an off switch
- Microsoft’s AI scraping memo points to a data supply chain problem
- Paint-Anything makes hex colors a first-class diffusion control
- Double descent as an implicit regularization story
- What Actually Makes a Tokeniser Good: Search Beats Objective
- Distillation needs calibration when the teacher is biased
- Bellman Policy Optimization cuts one moving part from RLVR
- Federated learning gets more practical when privacy and timing are treated together
- OptiFlow treats offline RL policy learning as sample matching
- Garry Tan’s distillation argument is really about AI capability access
- The useful part of calling Nvidia AI’s central bank
- CRISPR screens need learned experiment pickers, not bigger chatbots
- Recursive Self-Improvement Has a Roadmap Now, and Most of It Isn't Built
- Layer-selective unlearning aims at the parts of a model that remember
- On-policy distillation may be pruning tails, not teaching
- Optimizers Are Becoming Systems Choices, Not AdamW Replacements
- Reward choice changes how LLM forecasters are wrong
- Alignment baked into pretraining, not bolted on later
- Selective Trust: Why RAG Systems Should Learn When to Ignore Their Own Context
- MindForge trains coding agents on blank-repo software work
- Hyperball optimizers still need learning-rate discipline
- MIRROR trains vision models by making each modality teach the others
- OPD2 tries to distill reasoning by subtracting the base model
- RL post-training as procedure compression, not just skill amplification
- Timestamp drift is the quiet ASR failure that breaks real workflows
- DemoPSD Treats Teacher Disagreement as a Training Signal
- RLVR needs a taste model, not just a grader
- Synthetic QA Has a Selection Problem Before It Has a Training Problem
- Autodata Turns Synthetic Data Generation Into an Agent You Train
- Self-distillation can make models better on the first try and worse on the fifth