Search arXiv⌕ Search

arXiv · 2610.08321

Covariate-dependent Nonparametric $g$-modeling for regression via infinite Mixture-of-Expertizing class

Abstract

Empirical Bayes $g$-modeling captures unit-level heterogeneity by estimating a latent prior distribution from observed data. In the existing formulations, however, the prior is shared by all units. In this paper, we develop a covariate-dependent g-modeling framework for regression in which the entire prior distribution of the regression coefficients is allowed to depend on covariates. We formulate the estimation of the prior as nonparametric maximum likelihood estimation (NPMLE) of the covariate-dependent prior, and show that the unrestricted problem is ill-posed. To resolve this, we introduce the infinite Mixture-of-Expertizing class of conditional priors, under which the NPMLE is precisely a softmax-gated Mixture of Experts (MoE) whose number of experts is not fixed in advance but is determined by the data. Building on a first-order optimality condition, we propose two exemplar-based estimation algorithms that select experts automatically, together with a post-hoc aggregation of experts for interpretation. On the theoretical side, we show that every conditional prior in the class is Lipschitz continuous in the covariates, and that aggregated softmax gates can approximate any continuous gate function. The effectiveness of the proposed NPMLE is shown through application to synthetic datasets and real datasets.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Akira Okazaki, Keisuke Yano. 2026-10-06. Covariate-dependent Nonparametric $g$-modeling for regression via infinite Mixture-of-Expertizing class. https://arxiv.org/abs/2610.08321

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

On the permutation equivariance principle for causal estimands

In many causal inference problems, multiple action variables share a common causal role yet lack a natural ordering. \revblue{We consider $K\geq2$ action variables, each evaluated under treatment or control conditions,} and formalize permutation equivariance, the principle that permuting the variables permutes the corresponding estimands in a trackable manner, hence preserving their scientific meaning. We characterize this principle algebraically and present a complete class of weighted permutation equivariant estimands capturing main effects and interactions of all orders. We discuss the interpretation and choice of weights and characterize residual-free estimands, whose inclusion--exclusion sum recovers the endpoint contrast between the all-treated and all-control configurations. \revblue{Applying our general framework to network interference yields a new hierarchy of direct, indirect, and overall effects, whose first-order aggregates recover the average effects of \citet{hu2022average}. We also identify the overall effects as mixed derivatives of expected welfare under independent Bernoulli assignment.} We illustrate the framework through the contexts of factorial studies, causal mediation, and network interference.

stat.ME↗

Distinguishable Category Groupings in Complex Surveys: Which Categories Can the Design Tell Apart?

Surveys often estimate the share of cases in each of several categories, such as the causes of an event or the reasons for a visit. Readers then compare the categories, but reliability rules usually check each estimate alone, not the differences between them. We ask instead which groups of categories can be told apart under the sampling design. The categories often fall into blocks, such as driver-related and vehicle-related causes. A grouping is called distinguishable when every pair of its groups in the same block differs by more than the design's minimum detectable difference, computed from the primary sampling units at the design's degrees of freedom. The maximal distinguishable groupings are the most detailed and can differ in size, while merging two groups can destroy distinguishability. So even if splitting any one group in two breaks distinguishability, a more detailed distinguishable grouping may still exist. We propose two methods. One merges categories until the grouping is distinguishable, then tests every single split. The other lists every distinguishable grouping of a block with few categories. When groups in different blocks must also differ, there might be no grouping that qualifies. An application to the National Motor Vehicle Crash Causation Survey, with 24 primary sampling units in 12 strata, gives the grouping maxima for two blocks of its categories. Of more than 27 million groupings of its 13 vehicle-related critical reasons, none with more than four groups is distinguishable, and its environment-related reasons resolve into at most four groups. Merging two groups of a distinguishable vehicle grouping destroys distinguishability in 5,239 of 31,773 cases. If groups from different blocks must also be told apart, no grouping qualifies. Analysts can apply these methods before publishing categorical comparisons from designs with few degrees of freedom.

stat.ME↗

Selection of functional predictors and smooth coefficient estimation for scalar-on-function regression models

In the framework of scalar-on-function regression models, in which several functional variables are employed to predict a scalar response, we propose a methodology for selecting relevant functional predictors while simultaneously providing accurate smooth (or, more generally, regular) estimates of the functional coefficients. We suppose that the functional predictors belong to a real separable Hilbert space, while the functional coefficients belong to a specific subspace of this Hilbert space. Such a subspace can be a Reproducing Kernel Hilbert Space (RKHS) to ensure the desired regularity characteristics, such as smoothness or periodicity, for the coefficient estimates. Our procedure, called SOFIA (Scalar-On-Function Integrated Adaptive Lasso), is based on an adaptive penalized least squares algorithm that leverages functional subgradients to efficiently solve the minimization problem. We demonstrate that the proposed method satisfies the functional oracle property, even when the number of predictors exceeds the sample size. SOFIA's effectiveness in variable selection and coefficient estimation is evaluated through extensive simulation studies and a real-data application to GDP growth prediction.

stat.ME↗