arXiv · 2609.06406
Hierarchical Wasserstein Merging for Multi-Domain Multi-Task Learning: From Specialists to a Generalist
Abstract
Multi-domain multi-task learning (MD-MTL) aims to build a single generalist model that performs well across heterogeneous domains and tasks. However, joint training often suffers from interference under distribution shifts. Existing model merging methods mostly operate on model parameters while overlooking the geometric structure of latent representation distributions across domains and tasks. To address these limitations, we propose Hierarchical Wasserstein Merging (HWM), a representation-level framework that models each domain-task specialist as a distribution of hidden representations on a shared support. HWM constructs task-level and global Wasserstein barycenters to capture within-task domain variation and cross-task structure, enabling either training-free specialist aggregation by Wasserstein-derived weights or training-based generalist learning through a hybrid Wasserstein alignment loss. Experiments on four NLP tasks across four domains per task show that HWM achieves superior effectiveness and generalization capability in MD-MTL settings.
Explore related subjects
Keep this discovery
Ming Cheng, Jiaying Gong, Hoda Eldardiry. 2026-09-06. Hierarchical Wasserstein Merging for Multi-Domain Multi-Task Learning: From Specialists to a Generalist. https://arxiv.org/abs/2609.06406
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.