Search arXivSearch

arXiv subjects

Masataka Taguri

Publications and source records attributed to Masataka Taguri.

9 recordsLinked to original sources

Efficient estimation of weighted treatment effects under two-phase sampling

Two-phase sampling offers a practical way to collect costly confounders only in a subsample while retaining inexpensive information for a larger cohort. In observational causal studies, however, phase-2 selection can distort estimation of population causal effects if the sampling mechanism is ignored, and available phase-1 information may also be exploited to improve efficiency. Yet efficiency theory for causal estimands under such designs remains limited, particularly beyond the average treatment effect. In this paper, we derive the semiparametric efficiency bound for a class of propensity-score-weighted average treatment effects, which includes the average treatment effect, effects among treated and untreated populations, and the overlap effect, under two-phase sampling. In addition to straightforward weighting estimators based on the known sampling probabilities, we propose an enriched doubly robust estimator that attains the efficiency bound when all nuisance functions are consistently estimated. In particular, under outcome-dependent sampling, substantial efficiency gains can arise in some settings by appropriately incorporating phase-1 information. We further conduct extensive simulation studies, varying the choice of phase-1 variables and sampling schemes, to characterize when and to what extent leveraging phase-1 information leads to efficiency gains.

stat.ME

On the uncertainty from the first-stage estimation of prognostic covariate adjustment in randomized controlled trials

Prognostic covariate adjustment (PROCOVA) is a two-sample two-stage estimation method for covariate adjustment in randomized controlled trials. In the first stage, a prognostic score, defined as the conditional expectation of an outcome given covariates under the control treatment, is estimated using historical data. In the second stage, analysis of covariance with the estimated prognostic score and treatment assignment as explanatory variables is performed, and the average treatment effect is estimated. Although the prognostic score is estimated in this procedure, the variance estimator, which treats the prognostic score as known, has been used. Furthermore, the difference in the asymptotic variance between cases where the prognostic score is known versus where it is estimated has not been previously clarified. In this study, we derived these two asymptotic variances and showed that they are equal. This result also holds when the prognostic score is estimated using machine learning with $L_2$ consistency. We also constructed two variance estimators: one that treats the prognostic score as known, and another that accounts for its estimation, and compared their performance through simulation studies and data applications. For PROCOVA, since both variance estimators are asymptotically valid, it is generally recommended to use a variance estimator that treats the prognostic score as known, as it is simpler to derive and implement. When historical data is small, a variance estimator that explicitly accounts for prognostic score estimation is recommended if conservative inference is preferred.

stat.ME

A Direct Variance Estimation (DiVE) for Meta-Analysis of Median Differences

Meta-analyses of two-group studies that report median differences typically rely on methods that require, in addition to the median difference and sample size, summary measures of dispersion such as quartiles or ranges. Studies that do not report such statistics are often excluded from the meta-analysis. Existing two-stage approaches first estimate the asymptotic variance of the median difference within each study under parametric assumptions, and then combine these study-specific estimates to obtain the pooled median difference and its variance. We propose Direct Variance Estimation (DiVE), a method that directly estimates the variance of the pooled difference using only study-level median differences and their sample sizes. A comprehensive simulation study across a wide range of distributional scenarios shows that DiVE performs comparably to or better than conventional two-stage methods, with clear advantages when the number of studies is small. A re-analysis of published meta-analyses demonstrates that DiVE enables the inclusion of studies lacking dispersion statistics, leading to a more comprehensive and potentially less biased synthesis of evidence.

stat.ME

On the Conservativeness of Robust Variance Estimators in Propensity Score Weighted Cox Models

In propensity score weighted analysis, robust variance that does not account for weight estimation is commonly used. In propensity score weighted Cox models (CoxPSW), the robust variance is known to be conservative when weights for the average treatment effect (ATE) are used, but it remains unclear whether this conservativeness also holds for other weighting schemes. This study evaluated the performance of the robust variance in CoxPSW when weights other than ATE are applied. We conducted an asymptotic comparison between the robust variance and a variance estimator that accounts for weight estimation under non-ATE weights. Their performance was further evaluated through simulation studies and real data analysis. The analytical results, simulations, and real data analysis indicated that the robust variance is not necessarily conservative in CoxPSW when weights other than ATE are used. These findings suggest that variance estimators that account for weight estimation should be used when applying non-ATE weights in CoxPSW.

stat.ME

Estimation of time-varying treatment effects using marginal structural models dependent on partial treatment history

Inverse probability (IP) weighting of marginal structural models (MSMs) can provide consistent estimators of time-varying treatment effects under correct model specifications and identifiability assumptions, even in the presence of time-varying confounding. However, this method has two problems: (i) inefficiency due to IP-weights cumulating all time points and (ii) bias and inefficiency due to the MSM misspecification. To address these problems, we propose (i) new IP-weights for estimating parameters of the MSM that depends on partial treatment history and (ii) closed testing procedures for selecting partial treatment history (how far back in time the MSM depends on past treatments). We derive the theoretical properties of our proposed methods under known IP-weights and discuss their extension to estimated IP-weights. Although some of our theoretical results are derived under additional assumptions beyond standard identifiability assumptions, some of which can be checked empirically from the data. In simulation studies, our proposed methods outperformed existing methods both in terms of performance in estimating time-varying treatment effects and in selecting partial treatment history. Our proposed methods have also been applied to real data of hemodialysis patients with reasonable results.

stat.ME

Simultaneous Modeling of Disease Screening and Severity Prediction: A Multi-task and Sparse Regularization Approach

Identifying clinically relevant biomarkers and developing predictive models are central challenges in biomedical research. Biomarkers are commonly used for disease screening, and some provide information not only on the presence or absence of a disease but also on its severity. Such biomarkers can contribute to treatment prioritization and support clinical decision-making. To address both disease screening and severity prediction, this paper focuses on regression modeling for ordinal outcomes with a hierarchical structure. When the response variable is a combination of the presence of disease and severity, such as {healthy, mild, intermediate, severe}, a straightforward approach is to apply the conventional ordinal regression model. However, such models may lack the flexibility needed to capture heterogeneity in how predictors relate to response levels, particularly when the response levels have a heterogeneous association structure with predictors. Therefore, this paper proposes a model that treats screening and severity prediction as separate tasks, along with an estimation method based on structural sparse regularization. This method is designed to leverage a shared structure between the tasks. In numerical experiments, the proposed method demonstrated stable performance across many scenarios compared to existing ordinal regression methods.

stat.ME

Robust Estimation and Model Selection for the Controlled Directed Effect with Unmeasured Mediator-Outcome Confounders

Controlled Direct Effect (CDE) is one of the causal estimands used to evaluate both exposure and mediation effects on an outcome. When there are unmeasured confounders existing between the mediator and the outcome, the ordinary identification assumption does not work. In this manuscript, we consider an identification condition to identify CDE in the presence of unmeasured confounders. The key assumptions are: 1) the random allocation of the exposure, and 2) the existence of instrumental variables directly related to the mediator. Under these conditions, we propose a novel doubly robust estimation method, which work well if either the propensity score model or the baseline outcome model is correctly specified. Additionally, we propose a Generalized Information Criterion (GIC)-based model selection criterion for CDE that ensures model selection consistency. Our proposed procedure and related methods are applied to both simulation and real datasets to confirm the performance of these methods. Our proposed method can select the correct model with high probability and accurately estimate CDE.

stat.ME

Nonparametric Bayesian Adjustment of Unmeasured Confounders in Cox Proportional Hazards Models

In observational studies, unmeasured confounders present a crucial challenge in accurately estimating desired causal effects. To calculate the hazard ratio (HR) in Cox proportional hazard models for time-to-event outcomes, two-stage residual inclusion and limited information maximum likelihood are typically employed. However, these methods are known to entail difficulty in terms of potential bias of HR estimates and parameter identification. This study introduces a novel nonparametric Bayesian method designed to estimate an unbiased HR, addressing concerns that previous research methods have had. Our proposed method consists of two phases: 1) detecting clusters based on the likelihood of the exposure and outcome variables, and 2) estimating the hazard ratio within each cluster. Although it is implicitly assumed that unmeasured confounders affect outcomes through cluster effects, our algorithm is well-suited for such data structures. The proposed Bayesian estimator has good performance compared with some competitors.

stat.ME

False Discovery Rate Control for Confounder Selection Using Mirror Statistics

While data-driven confounder selection requires careful consideration, it is frequently employed in observational studies. Widely recognized criteria for confounder selection include the minimal-set approach, which involves selecting variables relevant to both treatment and outcome, and the union-set approach, which involves selecting variables associated with either treatment or outcome. These approaches are often implemented using heuristics and off-the-shelf statistical methods, where the degree of uncertainty may not be clear. In this paper, we focus on the false discovery rate (FDR) to measure uncertainty in confounder selection. We define the FDR specific to confounder selection and propose methods based on the mirror statistic, a recently developed approach for FDR control that does not rely on p-values. The proposed methods are p-value-free and require only the assumption of some symmetry in the distribution of the mirror statistic. It can be combined with sparse estimation and other methods that involve difficulties in deriving p-values. The properties of the proposed methods are investigated through exhaustive numerical experiments. Particularly in high-dimensional data scenarios, the proposed methods effectively control FDR and perform better than the p-value-based methods.

stat.ME