Search arXiv⌕ Search

arXiv subjects

Kwangmoon Park

Publications and source records attributed to Kwangmoon Park.

4 recordsLinked to original sources

Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions

Causal discovery from observational and interventional data becomes challenging in the presence of latent confounding and selection bias, where causal structure is no longer adequately represented by directed acyclic graphs over observed variables. Existing model-free methods often rely on an exponential number of conditional independence tests and provide limited uncertainty quantification in high-dimensional settings. We develop a model-free and constraint-query optimal statistical inference framework for causal discovery under latent variables and selection using single-target interventions. We introduce the system-induced subgraph (SIS) to capture the causal relations among system variables while accounting for context variables. We establish its identifiability through maximal ancestral graphs (MAGs), and show that interventions on each observed system variable are sufficient for unique identification and necessary in the worst case. Building on these results, we develop a two-stage graph inference procedure with asymptotic family-wise error control under sufficient first-stage power. For $d_X$ observed system variables, the procedure requires at most $\frac{5}{2}d_X^2$ statistical tests, parallelizable within each stage, and achieves optimal constraint-query complexity up to a constant factor. The framework accommodates soft interventions and avoids parametric structural equation assumptions. We illustrate the methods through analysis of Perturb-seq data from interferon-$β$-stimulated A549 lung cancer cell lines.

stat.ME↗

Causal Path Analysis from Perturbational and Population-Scale Single-Cell Data with Multiscale Confounding and Measurement Error

Single-cell perturbation experiments provide causal information on gene regulation, whereas population-scale single-cell studies characterize gene expression and phenotypes in human populations. We develop a framework that integrates these complementary data sources for causal path analysis. Rather than assuming that a perturbational gene network transfers directly to the target population, we use externally learned ancestral relationships to constrain the network topology and re-estimate its direct edges and effects from population data. To address latent heterogeneity and measurement error in multiscale single-cell measurements, we develop a surrogate-variable procedure operating at both the cell and subject levels, combined with errors-in-variables correction for network and outcome regressions. We establish theoretical guarantees for confounder recovery and high-dimensional estimation of network and gene-outcome effects. Simulations demonstrate the importance of jointly correcting confounding and measurement error. An application to acute myeloid leukemia identifies distinct regulatory pathways linking transcriptional regulators to blast count.

stat.ME↗

Confounder-robust causal discovery and inference in Perturb-seq using proxy and instrumental variables

Emerging single-cell technologies that combine CRISPR-based genetic perturbations with single-cell RNA sequencing, such as Perturb-seq, offer unprecedented opportunities to uncover cause-and-effect relationships among genes. Nonetheless, Perturb-seq experiments are subject to unobserved factors that, if not properly handled, can severely bias the inferred causal relationships between genes. These latent factors may arise not only from intrinsic molecular features of the regulatory elements, but also from unmeasured genes omitted due to cost-constrained experimental designs. Although methods for analyzing large-scale Perturb-seq data are rapidly maturing, approaches that explicitly account for such unobserved confounders when inferring causal gene networks are still lacking. Here, we propose a novel approach to accurately reconstruct causal gene networks from Perturb-seq data even when important confounders are missing. Our framework leverages proxy and instrumental variable strategies to exploit the rich information embedded in the perturbations, enabling unbiased estimation of the underlying directed acyclic graph (DAG) of gene expression. Applications to both comprehensive synthetic data and real CRISPR interference experiments in K562 cells demonstrate that our method outperforms baseline approaches that lack principled adjustments for unmeasured confounding, yielding more accurate and biologically relevant recovery of the true causal DAGs.

stat.ME↗

Sparse higher order partial least squares for simultaneous variable selection, dimension reduction, and tensor denoising

Partial Least Squares (PLS) regression emerged as an alternative to ordinary least squares for addressing multicollinearity in a wide range of scientific applications. As multidimensional tensor data is becoming more widespread, tensor adaptations of PLS have been developed. In this paper, we first establish the statistical behavior of Higher Order PLS (HOPLS) of Zhao et al. (2012), by showing that the consistency of the HOPLS estimator cannot be guaranteed as the tensor dimensions and the number of features increase faster than the sample size. To tackle this issue, we propose Sparse Higher Order Partial Least Squares (SHOPS) regression and an accompanying algorithm. SHOPS simultaneously accommodates variable selection, dimension reduction, and tensor response denoising. We further establish the asymptotic results of the SHOPS algorithm under a high-dimensional regime. The results also complete the unknown theoretic properties of SPLS algorithm (Chun and Keleş, 2010). We verify these findings through comprehensive simulation experiments, and application to an emerging high-dimensional biological data analysis.

stat.ME↗