model-evaluation
12 posts tagged model-evaluation.
- Pachocki’s warning points to safety gates, not slower vibes
- Uncensored Qwen edits show why model cards are not enough
- Certified world models still have blind topology
- CAST makes clinical model audits more concrete
- Qwen thinking levels need task-level tests, not vibes
- TokEval makes tokenizer choice measurable before pretraining
- EPC scores explanations by testing what the model can lose
- Instruction following is the local model test benchmarks miss
- Vacuum 16T turns model size into a metadata bug
- Tabular foundation models still stumble when the rows change
- Uncertainty metrics should follow the loss, not the other way around
- New model releases do not reset the advantage