Search arXiv⌕ Search

arXiv · 2610.04135

Minimax Rates for Learning Smooth Populations of Parameters

Abstract

We study pointwise estimation of a smooth binomial mixing density and characterize its minimax rates. Assuming that the mixing density is $s$-Hölder smooth, we first consider the homogeneous setting with a common number of binomial trials $t$. We derive matching lower and upper bounds that reveal three regimes depending on the relative sizes of the sample size $n$ and the number of trials $t$. When $t$ is small, the finite number of trials creates an identification barrier that persists regardless of sample size; at intermediate $t$, more data reduce sampling uncertainty, while the number of trials limits how finely the mixing density can be recovered; and when $t$ is sufficiently large, the usual nonparametric density estimation rate is recovered. Our lower bounds exploit the polynomial structure of the binomial mixture model, while attainability is achieved by a suitably regularized orthogonal series estimator. We then extend the analysis to heterogeneous trial parameters, where the minimax rate depends on the full trial profile through the effective sample size available at each polynomial degree. These results characterize the fundamental statistical limits of learning smooth populations of binomial probabilities under both homogeneous and heterogeneous trials.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

JungHo Lee, Edward H. Kennedy. 2026-10-02. Minimax Rates for Learning Smooth Populations of Parameters. https://arxiv.org/abs/2610.04135

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Structural Causal Models for Extremes: an Approach Based on Exponent Measures

We introduce a new formulation of structural causal models for extremes, called the extremal structural causal model (eSCM). Unlike conventional structural causal models, where randomness is governed by a probability distribution, eSCMs use an exponent measure, an infinite-mass law that naturally arises in the analysis of multivariate extremes. Central to this framework are activation variables, which abstract the single-big-jump principle, along with additional randomization that enriches the class of eSCM laws. The eSCM provides a common foundation for the two main existing approaches to extremal causal modeling, the max-linear and the sum-linear structural causal models, recovering both as special cases within a single asymptotic formulation. More broadly, it encompasses all possible laws of directed graphical models under the recently introduced notion of extremal conditional independence. We also identify an inherent asymmetry in eSCMs under natural assumptions, enabling the identifiability of causal directions, a central challenge in causal inference. Finally, we propose a method that exploits this causal asymmetry and demonstrate its effectiveness on both simulated and real datasets.

math.ST↗

Consistency of penalized maximum likelihood estimation for multidimensional change-in-velocity detection

We establish statistical consistency for the penalized maximum-likelihood estimator underlying CPLASS, a method for detecting changes in velocity in $d$-dimensional time-series data. The signal is modeled as a continuous piecewise-linear trajectory observed with independent Gaussian noise. Unlike classical change-in-mean models, continuity across adjacent segments couples their parameters and prevents direct application of standard segmentation arguments. Under compact parameter spaces, minimum segment-length and velocity-jump conditions, and a strengthened Schwarz information criterion penalty $ρ_k(\log n)^γ$ with $γ>1$, we prove joint consistency of the estimated number of segments and all changepoint locations in the high-frequency regime $n\to\infty$. The maximum changepoint-location error is $O_{\mathbb{P}}\{(\log n/n)^{1/2}\}$. The proof first uses empirical-process theory to establish convergence of the fitted signal, variance, and likelihood for each fixed model size. It then combines overfitting and underfitting arguments with a two-stage geometric localization analysis, yielding an initial $(\log n/n)^{1/3}$ rate that is sharpened by exploiting the local two-segment continuous piecewise-linear structure.

math.ST↗

Optimal Community Recovery by Spectrally Initialized Variational EM in General Stochastic Block Models

We prove a Chernoff-exponent guarantee for the output of spectrally initialized batch variational EM after a prescribed iteration budget, without assuming global maximization of its variational objective. The iteration repeatedly estimates the entire block probability matrix and community proportions from the same sparse graph. We consider a fixed number of communities with proportions bounded away from zero and fixed, positive, distinct connectivity profiles; neither assortativity nor full rank is required. The algorithm uses simultaneous softmax updates without sample splitting or posterior thresholding. Its analysis must control the feedback between estimated parameters, soft labels and reused edges at an exponentially small risk scale. We establish a uniform one-step bound over data-dependent soft assignments whose random remainder has exponentially small expectation. Combined with an exponentially reliable regularized spectral initializer, this yields an unconditional expected misclassification rate bounded by $\exp\{-(1-o(1))J_n\}$ throughout the sparse, diverging-degree regime, where $J_n$ is the minimum nodewise Chernoff information. Matching lower bounds establish first-order logarithmic minimax optimality on local parameter spaces allowing unknown connectivity and varying community counts. The algorithm also attains the sharp first-order exact-recovery threshold. Numerical experiments illustrate refinement gains and sensitivity to initialization and imbalance.

math.ST↗