Search arXivSearch

arXiv · 2609.06098

Causal DAG Identification for Count Data via Poisson Thinning Structural Equation Models

Abstract

Count-valued variables arise in many scientific and applied settings, yet explicit structural models that allow full identification of causal DAGs from observational data remain limited. The Poisson branching structural causal model (PB-SCM) provides a count-valued analogue of linear structural equation models using binomial thinning and independent Poisson exogenous variables, but its causal DAG is generally only partially identifiable. Building on this framework, we propose the Poisson thinning structural equation model (PT-SEM), which replaces binomial thinning in PB-SCM with Poisson thinning and allows node-wise exogenous distributions from diverse count-distribution families. Under node-wise regularity conditions, we establish identifiability of the causal DAG, the thinning coefficients, and the node-wise exogenous distributions. The same identification analysis extends to binomial thinning, yielding full identifiability whenever every nonsink has non-Poisson exogenous noise. We further develop a structure learning algorithm that optimizes, via dynamic programming, a BIC score based on local likelihoods evaluated at plug-in moment estimates, and establish its consistency for DAG selection. Simulations demonstrate favorable performance in DAG recovery and thinning-coefficient estimation, and a real-data application illustrates the practical utility of PT-SEM.

Explore related subjects

Keep this discovery

BibTeXRIS

Penggang Gao, Ming Cai, Hisayuki Hara. 2026-09-05. Causal DAG Identification for Count Data via Poisson Thinning Structural Equation Models. https://arxiv.org/abs/2609.06098

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Clustering Three-Way Data with Outliers

Matrix-variate distributions are a relatively recent addition to the model-based clustering literature, thereby making it possible to analyze data in matrix form with complex structure such as images and time series. Due to its recent appearance, there is limited literature on matrix-variate data, with even less on dealing with outliers in these models. An approach for clustering matrix-variate normal data with outliers is discussed. The approach, which uses the distribution of subset log-likelihoods, extends the OCLUST algorithm to matrix-variate normal data and uses an iterative approach to detect and trim outliers.

stat.ML

Robust conditional dimension reduction for dissimilarity data

Conditional dimension reduction (cDR) learns low-dimensional latent coordinates while accounting for observed covariates that represent known sources of variation in the data. Conditional Multidimensional Scaling (cMDS) is a cDR technique that works directly with dissimilarity data. Its standard squared-stress formulation, however, is sensitive to contaminated dissimilarity, since outliers can dominate the objective and distort the learned configuration. We proposed Robust Conditional Multidimensional Scaling (rcMDS) by replacing the squared-stress criterion with a Fair M-estimation objective. We developed a reweighted conditional SMACOF algorithm to optimize this objective. The proposed algorithm admits computationally tractable updates, and its stabilized objective values decrease monotonically and converge to a finite limit. Experiments on synthetic and real data show that the pro

stat.ML

The Role of Uncertainty in Assessing the Fairness of Machine Learning Models

Machine learning models are widely used in clinical applications, social media, law enforcement and critical infrastructure. Verifying whether their outputs are biased against disadvantaged groups or individuals is crucial to ensuring they are fair and allowing their use in such settings. A rigorous risk assessment of possible fairness violations requires quantifying the uncertainty associated with selecting and estimating such models. Yet, this is rarely done in the literature, which focuses on identifying a single model with a suitable trade-off between predictive accuracy and fairness. In this paper, we move beyond point estimation and discuss frequentist and Bayesian approaches to uncertainty quantification for fair machine learning, with practical examples and implications for simulated and real data.

stat.ML