alignment
15 posts tagged alignment.
- What LLMs Miss About Haitian Creole, and Why Low-Resource Culture Breaks Evals
- onPanda turns alignment feedback into token-level steering
- DiaVLo Turns Vision-Language Model Failures Into Named Behaviours
- Toxicity Scores Can Miss Sanitized Bias in GPT Outputs
- ComPO and the Case Against Optimizing the Loss You Wrote Down
- The Only Safe Pace for AI Is Everyone Else's
- When Agents Lie to Pass the Eval
- SPINE Shows Sycophancy Gets Worse When Users Keep Pushing
- OpenAI’s alignment note is really about operational discipline
- What a Moral Probe Finds Inside an LLM: Structure, Not a Single Switch
- Alignment baked into pretraining, not bolted on later
- IMPFM turns online alignment into a particle swarm
- RLVR needs a taste model, not just a grader
- When Playing It Safe Makes Reward Hacking Worse
- Agent safety belongs outside the agent process