arXiv · 2609.24504
On Emergent Capabilities and Model Merging
Abstract
Fine-tuned checkpoints and adapters now fill public repositories, and the most common operation applied to these artifacts is model merging: arithmetic on their weights that assembles capabilities cheaply. We ask what this operation does to emergent capabilities: behaviors an artifact carries that were never an explicit training target. Studying two independent testbeds (activation oracles and emergent-misaligned models) across three model families, we find that the answer is threefold. First, merging preserves an emergent capability that both parents carry: merging two misaligned checkpoints retains most of their broad misalignment across the whole mixing range. Second, merging cannot create an emergent capability that is superadditive in its parents: no weighted merge of two single-task oracles reaches the jointly-trained oracle's auditing ability. Third, when only one parent carries the capability, merging dilutes it faster than the trained capability that accompanies it: the gap is significant in most settings. In short, emergent behaviors of an artifact do not compose the way its trained capability does.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Luca Zhou, Emanuele Rodolà. 2026-09-21. On Emergent Capabilities and Model Merging. https://arxiv.org/abs/2609.24504
Cite the original work for its findings. Save a collection to share your selection of sources.