arXiv · 2609.08966
Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
Abstract
Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training. We show that this assumption can fail in a full 30B mixture-of-experts training pipeline. The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.
Explore related subjects
Keep this discovery
Sohir Maskey, Philipp Scholl, Jonas Knupp, Pit Neitemeier, Sascha Wirges. 2026-09-08. Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack. https://arxiv.org/abs/2609.08966
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.