Search arXivSearch

arXiv subjects

Ming Cai

Publications and source records attributed to Ming Cai.

2 recordsLinked to original sources

Causal DAG Identification for Count Data via Poisson Thinning Structural Equation Models

Count-valued variables arise in many scientific and applied settings, yet explicit structural models that allow full identification of causal DAGs from observational data remain limited. The Poisson branching structural causal model (PB-SCM) provides a count-valued analogue of linear structural equation models using binomial thinning and independent Poisson exogenous variables, but its causal DAG is generally only partially identifiable. Building on this framework, we propose the Poisson thinning structural equation model (PT-SEM), which replaces binomial thinning in PB-SCM with Poisson thinning and allows node-wise exogenous distributions from diverse count-distribution families. Under node-wise regularity conditions, we establish identifiability of the causal DAG, the thinning coefficients, and the node-wise exogenous distributions. The same identification analysis extends to binomial thinning, yielding full identifiability whenever every nonsink has non-Poisson exogenous noise. We further develop a structure learning algorithm that optimizes, via dynamic programming, a BIC score based on local likelihoods evaluated at plug-in moment estimates, and establish its consistency for DAG selection. Simulations demonstrate favorable performance in DAG recovery and thinning-coefficient estimation, and a real-data application illustrates the practical utility of PT-SEM.

stat.ME

When Less is More: Understanding When Token Filtering Helps and Fails in AI-generated Text Detection

The rapid advancement of large language models (LLMs) has made AI-generated text detection increasingly critical. Existing zero-shot detectors assume that more token-level evidence leads to more reliable detection. However, our empirical study challenges this consensus: fewer tokens sometimes work better, retaining only 40% can yield optimal performance, yet this benefit is not universal. Using the Entropy Gap Score (EGS), we introduce top-$k$ cumulative probability filtering as a diagnostic probe. Across three representative settings, filtering exhibits strikingly different behaviors. We analyze EGS via typical set theory and quantify its dynamics through entropy calibration and distribution analysis. We find that filtering helps for weak source LMs, where low-entropy tokens are harmful, but fails for strong source LMs, where they are not notably harmful. Our work provides the first systematic analysis showing that some tokens are not merely uninformative but systematically harmful due to entropy miscalibration, revealing a two-sided trade-off in token-level detection.

cs.CL