Search arXiv⌕ Search

arXiv · 2610.08092

What Should We Measure Next? Finding Identification Strategies by Refining Mechanisms

Abstract

Canonical approaches in causal inference treat model specification as fixed, assuming that researchers directly translate all relevant domain knowledge into a causal model, which can then be used to deduce its logical implications. Yet, in practical applications, model specification is often an iterative process, and involves exploration and introspection: of the many aspects of the phenomenon one could investigate, which ones actually matter for the identification of the causal effect of interest? In this paper we study the problem of iterative identification in partially specified causal models. We focus on determining where observing variables that intercept a direct effect or a confounding path between two variables could enable identification in semi-Markovian models. We give necessary and sufficient graphical conditions for when observing such variables can render an unidentifiable query identifiable, together with an efficient algorithm for locating all such opportunities in a given causal diagram. Our results can help analysts better navigate the model space by drawing attention to the parts of their substantive knowledge that could result in a successful identification strategy.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mikko Väänänen, Fanyu Cui, Carlos Cinelli, Santtu Tikka. 2026-10-06. What Should We Measure Next? Finding Identification Strategies by Refining Mechanisms. https://arxiv.org/abs/2610.08092

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

On the permutation equivariance principle for causal estimands

In many causal inference problems, multiple action variables share a common causal role yet lack a natural ordering. \revblue{We consider $K\geq2$ action variables, each evaluated under treatment or control conditions,} and formalize permutation equivariance, the principle that permuting the variables permutes the corresponding estimands in a trackable manner, hence preserving their scientific meaning. We characterize this principle algebraically and present a complete class of weighted permutation equivariant estimands capturing main effects and interactions of all orders. We discuss the interpretation and choice of weights and characterize residual-free estimands, whose inclusion--exclusion sum recovers the endpoint contrast between the all-treated and all-control configurations. \revblue{Applying our general framework to network interference yields a new hierarchy of direct, indirect, and overall effects, whose first-order aggregates recover the average effects of \citet{hu2022average}. We also identify the overall effects as mixed derivatives of expected welfare under independent Bernoulli assignment.} We illustrate the framework through the contexts of factorial studies, causal mediation, and network interference.

stat.ME↗

Distinguishable Category Groupings in Complex Surveys: Which Categories Can the Design Tell Apart?

Surveys often estimate the share of cases in each of several categories, such as the causes of an event or the reasons for a visit. Readers then compare the categories, but reliability rules usually check each estimate alone, not the differences between them. We ask instead which groups of categories can be told apart under the sampling design. The categories often fall into blocks, such as driver-related and vehicle-related causes. A grouping is called distinguishable when every pair of its groups in the same block differs by more than the design's minimum detectable difference, computed from the primary sampling units at the design's degrees of freedom. The maximal distinguishable groupings are the most detailed and can differ in size, while merging two groups can destroy distinguishability. So even if splitting any one group in two breaks distinguishability, a more detailed distinguishable grouping may still exist. We propose two methods. One merges categories until the grouping is distinguishable, then tests every single split. The other lists every distinguishable grouping of a block with few categories. When groups in different blocks must also differ, there might be no grouping that qualifies. An application to the National Motor Vehicle Crash Causation Survey, with 24 primary sampling units in 12 strata, gives the grouping maxima for two blocks of its categories. Of more than 27 million groupings of its 13 vehicle-related critical reasons, none with more than four groups is distinguishable, and its environment-related reasons resolve into at most four groups. Merging two groups of a distinguishable vehicle grouping destroys distinguishability in 5,239 of 31,773 cases. If groups from different blocks must also be told apart, no grouping qualifies. Analysts can apply these methods before publishing categorical comparisons from designs with few degrees of freedom.

stat.ME↗

Selection of functional predictors and smooth coefficient estimation for scalar-on-function regression models

In the framework of scalar-on-function regression models, in which several functional variables are employed to predict a scalar response, we propose a methodology for selecting relevant functional predictors while simultaneously providing accurate smooth (or, more generally, regular) estimates of the functional coefficients. We suppose that the functional predictors belong to a real separable Hilbert space, while the functional coefficients belong to a specific subspace of this Hilbert space. Such a subspace can be a Reproducing Kernel Hilbert Space (RKHS) to ensure the desired regularity characteristics, such as smoothness or periodicity, for the coefficient estimates. Our procedure, called SOFIA (Scalar-On-Function Integrated Adaptive Lasso), is based on an adaptive penalized least squares algorithm that leverages functional subgradients to efficiently solve the minimization problem. We demonstrate that the proposed method satisfies the functional oracle property, even when the number of predictors exceeds the sample size. SOFIA's effectiveness in variable selection and coefficient estimation is evaluated through extensive simulation studies and a real-data application to GDP growth prediction.

stat.ME↗