Search arXivSearch

arXiv subjects

Johan Lim

Publications and source records attributed to Johan Lim.

At least 19 recordsLinked to original sources

Bayesian Donor Set Selection in Synthetic Controls

The Synthetic Control Method (SCM) is a widely used approach for assessing the effects of interventions by constructing a synthetic counterfactual using a donor set of untreated units. However, the effectiveness of SCM heavily relies on the careful selection of an appropriate donor set. In this paper, we propose a Bayesian hierarchical model that performs donor set selection while preserving the standard SCM simplex constraint on donor weights. Unlike approaches that assume a fixed donor set, our model allows for the simultaneous estimation of the synthetic control weights and the active donor set. By using a hierarchical Gamma-Bernoulli construction for the donor weights, the proposed model assigns posterior mass to simplex faces and allows exact zero weights for excluded donors. We establish a posterior donor-set consistency result under a simplified pre-intervention model. Through numerical simulations, we show that our model improves donor recovery and weight estimation when the donor pool contains irrelevant or weakly related units, while remaining competitive in full-donor settings. Finally, we apply our model to the GDP trajectory of West Germany, illustrating its practical applicability. Our findings suggest that incorporating donor set selection offers a more parsimonious and flexible extension of existing Bayesian synthetic control methods.

stat.ME

Block-Independent Likelihood Ratio Testing for High-Dimensional Mean Vectors with Applications to Matrix-Variate Data

Testing the equality of two high-dimensional mean vectors is a fundamental problem in multivariate analysis. While the classical Hotelling's $T^2$ test is optimal in low-dimensional settings, it fails when the dimension $p$ is comparable to or exceeds the sample size $n$. Several extensions, including the Diagonal Likelihood Ratio Test (DLRT), have been proposed under the working independence assumption among variables. However, such an assumption can lead to a substantial loss of power when correlations are present. In this paper, we propose a new test, the Block Independent Likelihood Ratio Test (BILT), which generalizes DLRT by relaxing the working independence assumption to a block independence assumption. We establish its asymptotic normality of the null distribution of the BILT statistic for 'increasing $p$ with small $n$' under mild regularity conditions. We further analyze the asymptotic power of BILT under a local alternatives. Extensive simulation studies show that BILT maintains Type I error control and achieves substantially higher power than DLRT across a wide range of covariance structures. An application to the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset further demonstrates the application of BILT to testing mean differences between two matrix-variate populations.

stat.ME

Uncertainty-Aware Ideal Point Estimation via Variational EM

Roll-call data analysis aims to estimate legislators' ideal points and quantify the associated uncertainty. Existing approaches either rely on Bayesian methods implemented via Markov chain Monte Carlo sampling or focus primarily on point estimation, with uncertainty typically assessed through resampling procedures such as the bootstrap. Consequently, the computational burden of these approaches can become substantial when applied to large roll-call datasets. To address this challenge, we propose a computationally efficient likelihood method for estimating ideal points and their standard errors. Leveraging the P\'{o}lya--Gamma identity, we develop a variational expectation--maximization algorithm for estimating ideal points and introduce a variational Louis' method to approximate the observed Fisher information for standard error estimation. Numerical studies and applications to U.S. congressional roll-call data demonstrate that the proposed method produces accurate ideal point estimates and reliable standard errors while being substantially more computationally efficient than existing approaches.

stat.ME

Conformalized Method for Empirical Bayes Normal Mean Inference Problem with Heteroscedastic Variance

We study the normal mean inference problem, which involves simultaneous testing of the means of many normal distributions. This problem has been extensively studied within the empirical Bayes (EB) framework. However, the reliability of most EB methods heavily depends on two key conditions: (i) the prior distribution is correctly specified, and (ii) it can be accurately estimated. In practice, both conditions are difficult to satisfy, and it is often unclear whether they hold in a given application. To overcome these limitations, we propose a new algorithm, called COIN (COnformal Inference for Normal mean inference problem). Unlike traditional empirical Bayes approaches, COIN produces decision rules whose validity does not depend on the correct specification or accurate estimation of the prior. We theoretically prove that COIN asymptotically controls the false discovery rate at the nominal level, even in the presence of prior misspecification or estimation errors. Since the COIN algorithm requires an external training dataset to estimate the prior distribution and conformity score function, we introduce two data-splitting strategies -- sample-splitting and feature-splitting -- for the case where such external data are unavailable. We provide theoretical guarantees for the data-splitting strategies and demonstrate their effectiveness through extensive numerical studies and three real data examples.

stat.ME

On parameter estimation for the truncated skew-normal distribution

Parameter estimation for the truncated skew-normal distribution is challenging, as truncation introduces additional nonlinearity into the likelihood function and often leads to numerical instability in existing estimation procedures. In this paper, we propose a grid-based estimation method, referred to as GRID-MOM, for parameter estimation in the truncated skew-normal distribution. The proposed approach fixes the shape parameter on a pre-specified grid and, for each grid point, estimates the location and scale parameters using the method of moments. The optimal value of the shape parameter is then selected via likelihood-based comparison, yielding the final parameter estimates. By decoupling the estimation of the shape parameter from that of the location and scale parameters, the proposed method reduces the complexity of the optimization problem and improves numerical stability. We evaluate the finite-sample performance of the proposed estimator through an extensive numerical study, comparing it with existing methods under a variety of scenarios. The results demonstrate that the proposed method provides stable and accurate estimation, particularly for the shape parameter, suggesting that the proposed method offers a practical alternative for inference in truncated skew-normal models. We further demonstrate the practical applicability of the proposed method using phosphoproteomics data and hospital admission data.

stat.ME

Two-Stage Multiple Test Procedures Controlling False Discovery Rate with auxiliary variable and their Application to Set4Delta Mutant Data

In this paper, we present novel methodologies that incorporate auxiliary variables for multiple hypotheses testing related to the main point of interest while effectively controlling the false discovery rate. When dealing with multiple tests concerning the primary variable of interest, researchers can use auxiliary variables to set preconditions for the significance of primary variables, thereby enhancing test efficacy. Depending on the auxiliary variable's role, we propose two approaches: one terminates testing of the primary variable if it does not meet predefined conditions, and the other adjusts the evaluation criteria based on the auxiliary variable. Employing the copula method, we elucidate the dependence between the auxiliary and primary variables by deriving their joint distribution from individual marginal distributions.Our numerical studies, compared with existing methods, demonstrate that the proposed methodologies effectively control the FDR and yield greater statistical power than previous approaches solely based on the primary variable. As an illustrative example, we apply our methods to the Set4$\Delta$ mutant dataset. Our findings highlight the distinctions between our methodologies and traditional approaches, emphasising the potential advantages of our methods in introducing the auxiliary variable for selecting more genes.

stat.ME

Multiple Testing of One-Sided Hypotheses with Conservative $p$-values

We study a large-scale one-sided multiple testing problem in which test statistics follow normal distributions with unit variance, and the goal is to identify signals with positive mean effects. A conventional approach is to compute $p$-values under the assumption that all null means are exactly zero and then apply standard multiple testing procedures such as the Benjamini-Hochberg (BH) or Storey-BH method. However, because the null hypothesis is composite, some null means may be strictly negative. In this case, the resulting $p$-values are conservative, leading to a substantial loss of power. Existing methods address this issue by modifying the multiple testing procedure itself, for example through conditioning strategies or discarding rules. In contrast, we focus on correcting the $p$-values so that they are exact under the null. Specifically, we estimate the marginal null distribution of the test statistics within an empirical Bayes framework and construct refined $p$-values based on this estimated distribution. These refined $p$-values can then be directly used in standard multiple testing procedures without modification. Extensive simulation studies show that the proposed method substantially improves power when conventional $p$-values are conservative, while achieving comparable performance to existing methods when conventional $p$-values are exact. An application to phosphorylation data further demonstrates the practical effectiveness of our approach.

stat.ME

Empirical Bayes Method for Large Scale Multiple Testing with Heteroscedastic Errors

In this paper, we address the normal mean inference problem, which involves testing multiple means of normal random variables with heteroscedastic variances. Most existing empirical Bayes methods for this setting are developed under restrictive assumptions, such as the scaled inverse-chi-squared prior for variances and unimodality for the non-null mean distribution. However, when either of these assumptions is violated, these methods often fail to control the false discovery rate (FDR) at the target level or suffer from a substantial loss of power. To overcome these limitations, we propose a new empirical Bayes method, gg-Mix, which assumes only independence between the normal means and variances, without imposing any structural restrictions on their distributions. We thoroughly evaluate the FDR control and power of gg-Mix through extensive numerical studies and demonstrate its superior performance compared to existing methods. Finally, we apply gg-Mix to three real data examples to further illustrate the practical advantages of our approach.

stat.ME

$\ell_0$-Regularized Item Response Theory Model for Robust Ideal Point Estimation

Ideal point estimation methods face a significant challenge when legislators engage in protest voting -- strategically voting against their party to express dissatisfaction. Such votes introduce attenuation bias, making ideologically extreme legislators appear artificially moderate. We propose a novel statistical framework that extends the fast EM-based estimation approach of \cite{Imai2016} using $\ell_0$ regularization method to handle protest votes. Through simulation studies, we demonstrate that our proposed method maintains estimation accuracy even with high proportions of protest votes, while being substantially faster than MCMC-based methods. Applying our method to the 116th and 117th U.S. House of Representatives, we successfully recover the extreme liberal positions of ``the Squad'', whose protest votes had caused conventional methods to misclassify them as moderates. While conventional methods rank Ocasio-Cortez as more conservative than 69\% of Democrats, our method places her firmly in the progressive wing, aligning with her documented policy positions. This approach provides both robust ideal point estimates and systematic identification of protest votes, facilitating deeper analysis of strategic voting behavior in legislatures.

stat.AP

Graph Estimation Based on Neighborhood Selection for Matrix-variate Data

Undirected graphical models are powerful tools for uncovering complex relationships among high-dimensional variables. This paper aims to fully recover the structure of an undirected graphical model when the data naturally take matrix form, such as temporal multivariate data. As conventional vector-variate analyses have clear limitations in handling such matrix-structured data, several approaches have been proposed, mostly relying on the likelihood of the Gaussian distribution with a separable covariance structure. Although some of these methods provide theoretical guarantees against false inclusions (i.e. all identified edges exist in the true graph), they may suffer from crucial limitations: (1) failure to detect important true edges, or (2) dependency on conditions for the estimators that have not been verified. We propose a novel regression-based method for estimating matrix graphical models, based on the relationship between partial correlations and regression coefficients. Adopting the primal-dual witness technique from the regression framework, we derive a non-asymptotic inequality for exact recovery of an edge set. Under suitable regularity conditions, our method consistently identifies the true edge set with high probability. Through simulation studies, we compare the support recovery performance of the proposed method against existing alternatives. We also apply our method to an electroencephalography (EEG) dataset to estimate both the spatial brain network among 64 electrodes and the temporal network across 256 time points.

stat.ME

Strong Consistency of Sparse K-means Clustering

In this paper, we study the strong consistency of the sparse K-means clustering for high dimensional data. We prove the consistency in both risk and clustering for the Euclidean distance. We discuss the characterization of the limit of the clustering under some special cases. For the general (non-Euclidean) distance, we prove the consistency in risk. Our result naturally extends to other models with the same objective function but different constraints such as l0 or l1 penalty in recent literature.

math.ST

Linear Shrinkage Convexification of Penalized Linear Regression With Missing Data

One of the common challenges faced by researchers in recent data analysis is missing values. In the context of penalized linear regression, which has been extensively explored over several decades, missing values introduce bias and yield a non-positive definite covariance matrix of the covariates, rendering the least square loss function non-convex. In this paper, we propose a novel procedure called the linear shrinkage positive definite (LPD) modification to address this issue. The LPD modification aims to modify the covariance matrix of the covariates in order to ensure consistency and positive definiteness. Employing the new covariance estimator, we are able to transform the penalized regression problem into a convex one, thereby facilitating the identification of sparse solutions. Notably, the LPD modification is computationally efficient and can be expressed analytically. In the presence of missing values, we establish the selection consistency and prove the convergence rate of the $\ell_1$-penalized regression estimator with LPD, showing an $\ell_2$-error convergence rate of square-root of $\log p$ over $n$ by a factor of $(s_0)^{3/2}$ ($s_0$: the number of non-zero coefficients). To further evaluate the effectiveness of our approach, we analyze real data from the Genomics of Drug Sensitivity in Cancer (GDSC) dataset. This dataset provides incomplete measurements of drug sensitivities of cell lines and their protein expressions. We conduct a series of penalized linear regression models with each sensitivity value serving as a response variable and protein expressions as explanatory variables.

stat.ME

Sparse Hanson-Wright Inequality for a Bilinear Form of Sub-Gaussian Variables

In this paper, we derive a new version of Hanson-Wright inequality for a sparse bilinear form of sub-Gaussian variables. Our results are generalization of previous deviation inequalities that consider either sparse quadratic forms or dense bilinear forms. We apply the new concentration inequality to testing the cross-covariance matrix when data are subject to missing. Using our results, we can find a threshold value of correlations that controls the family-wise error rate. Furthermore, we discuss the multiplicative measurement error case for the bilinear form with a boundedness condition.

math.ST

Testing Independence of Bivariate Censored Data using Random Walk on Restricted Permutation Graph

In this paper, we propose a procedure to test the independence of bivariate censored data, which is generic and applicable to any censoring types in the literature. To test the hypothesis, we consider a rank-based statistic, Kendall's tau statistic. The censored data defines a restricted permutation space of all possible ranks of the observations. We propose the statistic, the average of Kendall's tau over the ranks in the restricted permutation space. To evaluate the statistic and its reference distribution, we develop a Markov chain Monte Carlo (MCMC) procedure to obtain uniform samples on the restricted permutation space and numerically approximate the null distribution of the averaged Kendall's tau. We apply the procedure to three real data examples with different censoring types, and compare the results with those by existing methods. We conclude the paper with some additional discussions not given in the main body of the paper.

stat.ME

Empirical Likelihood Inference for Area under the ROC Curve using Ranked Set Samples

The area under a receiver operating characteristic curve (AUC) is a useful tool to assess the performance of continuous-scale diagnostic tests on binary classification. In this article, we propose an empirical likelihood (EL) method to construct confidence intervals for the AUC from data collected by ranked set sampling (RSS). The proposed EL-based method enables inferences without assumptions required in existing nonparametric methods and takes advantage of the sampling efficiency of RSS. We show that for both balanced and unbalanced RSS, the EL-based point estimate is the Mann-Whitney statistic, and confidence intervals can be obtained from a scaled chi-square distribution. Simulation studies and two case studies on diabetes and chronic kidney disease data suggest that using the proposed method and RSS enables more efficient inference on the AUC.

stat.ME

Estimating High-dimensional Covariance and Precision Matrices under General Missing Dependence

A sample covariance matrix $\boldsymbol{S}$ of completely observed data is the key statistic in a large variety of multivariate statistical procedures, such as structured covariance/precision matrix estimation, principal component analysis, and testing of equality of mean vectors. However, when the data are partially observed, the sample covariance matrix from the available data is biased and does not provide valid multivariate procedures. To correct the bias, a simple adjustment method called inverse probability weighting (IPW) has been used in previous research, yielding the IPW estimator. The estimator plays the role of $\boldsymbol{S}$ in the missing data context so that it can be plugged into off-the-shelf multivariate procedures. However, theoretical properties (e.g. concentration) of the IPW estimator have been only established under very simple missing structures; every variable of each sample is independently subject to missing with equal probability. We investigate the deviation of the IPW estimator when observations are partially observed under general missing dependency. We prove the optimal convergence rate $O_p(\sqrt{\log p / n})$ of the IPW estimator based on the element-wise maximum norm. We also derive similar deviation results even when implicit assumptions (known mean and/or missing probability) are relaxed. The optimal rate is especially crucial in estimating a precision matrix, because of the "meta-theorem" that claims the rate of the IPW estimator governs that of the resulting precision matrix estimator. In the simulation study, we discuss non-positive semi-definiteness of the IPW estimator and compare the estimator with imputation methods, which are practically important.

stat.ME

Covariate-dependent control limits for the detection of abnormal price changes in scanner data

Currently, large-scale sales data for consumer goods, called scanner data, are obtained by scanning the bar codes of individual products at the points of sale of retail outlets. Many national statistical offices use scanner data to build consumer price statistics. In this process, as in other statistical procedures, the detection of abnormal transactions in sales prices is an important step in the analysis. Popular methods for conducting such outlier detection are the quartile method, the Hidiroglou-Berthelot method, the resistant fences method, and the Tukey algorithm. These methods are based solely on information about price changes and not on any of the other covariates (e.g., sales volume or types of retail shops) that are also available from scanner data. In this paper, we propose a new method to detect abnormal price changes that takes into account an additional covariate, namely, sales volume. We assume that the variance of the log of the price change is a smooth function of the sales volume and estimate the function from previously observed data. We numerically show the advantages of the new method over existing methods. We also apply the methods to real scanner data collected at weekly intervals by the Korean Chamber of Commerce and Industry between 2013 and 2014 and compare their performance.

stat.ME

Testing for Stochastic Order in Interval-Valued Data

We construct a procedure to test the stochastic order of two samples of interval-valued data. We propose a test statistic which belongs to U-statistic and derive its asymptotic distribution under the null hypothesis. We compare the performance of the newly proposed method with the existing one-sided bivariate Kolmogorov-Smirnov test using real data and simulated data.

stat.ME