Search arXivSearch

arXiv · 2505.24281

Multi-Task Learning with Covariate-Overlap Regularization

Abstract

Multi-task learning improves data efficiency by sharing information across related tasks, but indiscriminate sharing can be harmful when their covariate distributions and response relationships differ. We propose COVariate-ovERlap regularized multi-task learning (COVER) to address covariate and posterior heterogeneity. The model combines a common component function with a shared neural representation and low-dimensional task-specific coefficients. Taskwise second-moment matrices summarize covariate heterogeneity and determine the strength of coefficient integration in each representation direction. We derive a covariate-overlap penalty by minimizing the total squared change in two task predictors when their coefficients are replaced by one auxiliary coefficient. An equivalent auxiliary formulation supports end-to-end training without matrix inversion. An exact fixed-representation bias--variance decomposition quantifies how covariate overlap controls variance reduction and how posterior heterogeneity determines shrinkage bias. Global and localized end-to-end oracle inequalities account for jointly learning the neural functions and estimating the overlap matrices from the same observations. We give explicit neural-network rates and sharpen the stochastic prediction term when the regularized oracle risk and overlap-estimation error are small. Simulations across diverse heterogeneity settings show competitive performance against deep-learning and statistical data-integration methods, with the largest gains under joint heterogeneity. In a GTEx central-nervous-system analysis, COVER achieves the lowest response-averaged prediction error among the compared methods and reveals tissue-pair integration patterns.

Explore related subjects

Keep this discovery

BibTeXRIS

Yang Sui, Qi Xu, Yang Bai, Annie Qu. 2026-09-05. Multi-Task Learning with Covariate-Overlap Regularization. https://arxiv.org/abs/2505.24281

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Clustering Three-Way Data with Outliers

Matrix-variate distributions are a relatively recent addition to the model-based clustering literature, thereby making it possible to analyze data in matrix form with complex structure such as images and time series. Due to its recent appearance, there is limited literature on matrix-variate data, with even less on dealing with outliers in these models. An approach for clustering matrix-variate normal data with outliers is discussed. The approach, which uses the distribution of subset log-likelihoods, extends the OCLUST algorithm to matrix-variate normal data and uses an iterative approach to detect and trim outliers.

stat.ML

Robust conditional dimension reduction for dissimilarity data

Conditional dimension reduction (cDR) learns low-dimensional latent coordinates while accounting for observed covariates that represent known sources of variation in the data. Conditional Multidimensional Scaling (cMDS) is a cDR technique that works directly with dissimilarity data. Its standard squared-stress formulation, however, is sensitive to contaminated dissimilarity, since outliers can dominate the objective and distort the learned configuration. We proposed Robust Conditional Multidimensional Scaling (rcMDS) by replacing the squared-stress criterion with a Fair M-estimation objective. We developed a reweighted conditional SMACOF algorithm to optimize this objective. The proposed algorithm admits computationally tractable updates, and its stabilized objective values decrease monotonically and converge to a finite limit. Experiments on synthetic and real data show that the pro

stat.ML

The Role of Uncertainty in Assessing the Fairness of Machine Learning Models

Machine learning models are widely used in clinical applications, social media, law enforcement and critical infrastructure. Verifying whether their outputs are biased against disadvantaged groups or individuals is crucial to ensuring they are fair and allowing their use in such settings. A rigorous risk assessment of possible fairness violations requires quantifying the uncertainty associated with selecting and estimating such models. Yet, this is rarely done in the literature, which focuses on identifying a single model with a suitable trade-off between predictive accuracy and fairness. In this paper, we move beyond point estimation and discuss frequentist and Bayesian approaches to uncertainty quantification for fair machine learning, with practical examples and implications for simulated and real data.

stat.ML