Search arXivSearch

arXiv · 2504.08332

High-dimensional Clustering and Signal Recovery under Block Signals

Abstract

This paper studies computationally efficient methods and their minimax optimality for high-dimensional clustering and signal recovery under block signal structures. We propose two sets of methods, cross-block feature aggregation PCA (CFA-PCA) and moving average PCA (MA-PCA), designed for sparse and dense block signals, respectively. Both methods adaptively utilize block signal structures, applicable to non-Gaussian data with heterogeneous variances and non-diagonal covariance matrices. Specifically, the CFA method utilizes a block-wise U-statistic to aggregate and select block signals non-parametrically from data with unknown cluster labels. We show that the proposed methods are consistent for both clustering and signal recovery under mild conditions and weaker signal strengths than the existing methods without considering block structures of signals. Furthermore, we derive both statistical and computational minimax lower bounds (SMLB and CMLB) for high-dimensional clustering and signal recovery under block signals, where the CMLBs are restricted to algorithms with polynomial computation complexity. The minimax boundaries partition signals into regions of impossibility and possibility. No algorithm (or no polynomial time algorithm) can achieve consistent clustering or signal recovery if the signals fall into the statistical (or computational) region of impossibility. We show that the proposed CFA-PCA and MA-PCA methods can achieve the CMLBs for the sparse and dense block signal regimes, respectively, indicating the proposed methods are computationally minimax optimal. A tuning parameter selection method is proposed based on post-clustering signal recovery results. Simulation studies are conducted to evaluate the proposed methods. A case study on global temperature change demonstrates their utility in practice.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wu Su, Yumou Qiu. 2025-04-11. High-dimensional Clustering and Signal Recovery under Block Signals. https://arxiv.org/abs/2504.08332

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Bias-Correction for Privacy-Protected Spatial Autoregressive Models with Application to Restaurant Network Analysis

Spatial autoregressive (SAR) models and their extensions are important tools for studying network effects. However, with an increasing emphasis on data privacy, data providers often implement protection measures that render standard SAR models inapplicable. In this study, we introduce a privacy-protected SAR model that incorporates noise into both the response and covariates to meet privacy requirements. With noise present in both components, the traditional quasi-maximum likelihood estimator becomes difficult to compute because the likelihood function cannot be directly formulated. To bypass this hurdle, we begin with a pseudo-likelihood approach, initially omitting the noise in the covariates. A Newton-Raphson algorithm is then applied to compute the estimator; however, the estimator is biased. To address this, we propose a bias-corrected Newton-Raphson-type algorithm that simultaneously accounts for noise in both the response and covariates. We further show, under appropriate regularity conditions, that the resulting estimator is consistent and asymptotically normal. To further enhance computational efficiency, we also develop a bias-corrected least squares estimator. Several extensions are discussed, and the finite-sample performance of the proposed methods is evaluated through extensive simulations. We apply the proposed methodology to restaurant transaction data from a third-party payment platform. Our method identifies a statistically significant competitive network effect among restaurants and further reveals meaningful restaurant-customer interaction patterns.

stat.ME

A variational framework for modal estimation

Multivariate mode estimation arises in many statistical problems such as inverse problems, multimodal sampling, and density-based clustering, but becomes challenging in moderate to high dimensions, especially when the underlying density is not directly evaluable. We introduce GERVE (Gibbs-measure Entropy-Regularized Variational Estimation), a sample-based method for estimating multivariate modes by approximating Gibbs distributions directly from samples, without estimating or evaluating the density. GERVE uses Gaussian-mixture variational annealing and natural-gradient optimization, producing a mixture concentrated in high-density regions whose component responsibilities also provide a clustering of the observations. We prove theoretical guarantees in two regimes: as the Gibbs temperature goes to zero, the optimal variational mixture concentrates around the global modes of the population density; at fixed positive temperature, we prove existence, consistency, and asymptotic normality of empirical maximizers and propose a bootstrap procedure for uncertainty quantification. Simulations and a real-data experiment show that GERVE accurately recovers modes and produces meaningful clusters.

stat.ME

Objective Model Prior Probabilities in Variable Selection

For many years it was routine to use equal model prior probabilities in Bayesian model uncertainty analysis. At least twenty years ago it became clear that this was problematic, leading to support of much too large models in the increasingly huge model spaces being considered in genomics and other fields. A popular replacement was to adopt a suggestion of Harold Jeffreys for the variable selection problem in which a total of $k$ possible variables are being considered for inclusion in the model: give the collection of all models containing $d$ variables ($d = 0, . . . , k$) prior probability $1/(k + 1)$ and then divide this prior probability equally among the models in the collection. Many other choices of model prior probabilities that impose severe parsimony have also been introduced. We begin by reviewing the problems with using equal model prior probabilities and then discuss some serious problems with the Jeffreys choice. Finally, we introduce and study a number of objective alternative choices of model prior probabilities, from both numerical and theoretical perspectives.

stat.ME