Search arXivSearch

arXiv · 2505.15215

Clustering and Pruning in Causal Data Fusion

Abstract

Data fusion, the process of combining observational and experimental data, can enable the identification of causal effects that would otherwise remain non-identifiable. Although identification algorithms have been developed for specific scenarios, do-calculus remains the only general-purpose tool for causal data fusion, particularly when variables are present in some data sources but not others. However, approaches based on do-calculus may encounter computational challenges as the number of variables increases and the causal graph grows in complexity. Consequently, there exists a need to reduce the size of such models while preserving the essential features. For this purpose, we propose pruning (removing unnecessary variables) and clustering (combining variables) as preprocessing operations for causal data fusion. We generalize earlier results on a single data source and derive conditions for applying pruning and clustering in the case of multiple data sources. We give sufficient conditions for inferring the identifiability or non-identifiability of a causal effect in a larger graph based on a smaller graph and show how to obtain the corresponding identifying functional for identifiable causal effects. Examples from epidemiology and social science demonstrate the use of the results.

Explore related subjects

Keep this discovery

BibTeXRIS

Otto Tabell, Santtu Tikka, Juha Karvanen. 2026-09-08. Clustering and Pruning in Causal Data Fusion. https://arxiv.org/abs/2505.15215

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Clustering Three-Way Data with Outliers

Matrix-variate distributions are a relatively recent addition to the model-based clustering literature, thereby making it possible to analyze data in matrix form with complex structure such as images and time series. Due to its recent appearance, there is limited literature on matrix-variate data, with even less on dealing with outliers in these models. An approach for clustering matrix-variate normal data with outliers is discussed. The approach, which uses the distribution of subset log-likelihoods, extends the OCLUST algorithm to matrix-variate normal data and uses an iterative approach to detect and trim outliers.

stat.ML

Robust conditional dimension reduction for dissimilarity data

Conditional dimension reduction (cDR) learns low-dimensional latent coordinates while accounting for observed covariates that represent known sources of variation in the data. Conditional Multidimensional Scaling (cMDS) is a cDR technique that works directly with dissimilarity data. Its standard squared-stress formulation, however, is sensitive to contaminated dissimilarity, since outliers can dominate the objective and distort the learned configuration. We proposed Robust Conditional Multidimensional Scaling (rcMDS) by replacing the squared-stress criterion with a Fair M-estimation objective. We developed a reweighted conditional SMACOF algorithm to optimize this objective. The proposed algorithm admits computationally tractable updates, and its stabilized objective values decrease monotonically and converge to a finite limit. Experiments on synthetic and real data show that the pro

stat.ML

The Role of Uncertainty in Assessing the Fairness of Machine Learning Models

Machine learning models are widely used in clinical applications, social media, law enforcement and critical infrastructure. Verifying whether their outputs are biased against disadvantaged groups or individuals is crucial to ensuring they are fair and allowing their use in such settings. A rigorous risk assessment of possible fairness violations requires quantifying the uncertainty associated with selecting and estimating such models. Yet, this is rarely done in the literature, which focuses on identifying a single model with a suitable trade-off between predictive accuracy and fairness. In this paper, we move beyond point estimation and discuss frequentist and Bayesian approaches to uncertainty quantification for fair machine learning, with practical examples and implications for simulated and real data.

stat.ML