model-evaluation
6 posts tagged model-evaluation.
- EPC scores explanations by testing what the model can lose
- Instruction following is the local model test benchmarks miss
- Vacuum 16T turns model size into a metadata bug
- Tabular foundation models still stumble when the rows change
- Uncertainty metrics should follow the loss, not the other way around
- New model releases do not reset the advantage