Search arXiv⌕ Search

arXiv subjects

Qirui Hu

Publications and source records attributed to Qirui Hu.

At least 19 recordsLinked to original sources

Approximation Theorems for High-Dimensional Canonical U-Statistics: Gaussian Chaos and Phase Transition

We study simultaneous inference for maxima of canonical order-two $U$-statistics in high dimension. Degeneracy makes quadratic fluctuations leading, so ordinary Gaussian calibration can fail even after exact variance normalization. We show that the appropriate general target is a joint signed Gaussian quadratic chaos and establish a general approximation result that permits indefinite kernels. The general anti-concentration bound is too crude for high-dimensional inference, and we obtain sharper bounds under additional spectral structure. We also identify a phase transition from a non-Gaussian signed-chaos maximum to its covariance-matched Gaussian counterpart driven by the effective rank. For feasible inference, we propose a Gaussian multiplier bootstrap that avoid estimating eigensystems, and establish its validity. Two applications and extensive numerical simulations further illustrate the scope and practical performance of the proposed framework.

math.ST↗

Riemannian Simultaneous Inference for Tangent Vector Field Regression

We consider nonparametric tangent vector field regression on a Riemannian manifold without boundary. Because responses at different points lie in different tangent spaces, the proposed kernel estimator first parallel transports nearby responses to the target tangent space and then forms a volume-corrected local average. We first derive its uniform second-order bias, finite-bandwidth covariance, and stochastic rate. For simultaneous inference, the tangent norm is written as a supremum over the unit tangent bundle. Exact covariance whitening gives a unit-variance Gaussian field whose correlation length is of order $h$ along the base manifold and of order one along the fibre. Its local covariance geometry leads to a Gumbel limit with an explicit intrinsic constant. Combining this limit with Gaussian approximation and cross-fitted covariance estimation yields a feasible simultaneous confidence tube for the regression field. We further discuss improved finite-sample inference with bandwidth selection and high-order bias corrections. Simulations on various manifolds support the proposed inference procedure. A randomized reconstruction of global wind data illustrates how the tube's cross-sections describe spatially varying uncertainty.

stat.ML↗

Locally Private Inference for Riemannian Stochastic Optimization

We develop inference for manifold-valued population minimizers when each observation belongs to a different participant and only locally private messages reach the analyst. The method releases randomized tangent gradients and combines them through Riemannian stochastic approximation and Polyak-Ruppert averaging. Directly inserting a private data surrogate into a nonlinear loss can shift its population target, whereas conditional centring of the released gradient preserves the first-order equation. We introduce symmetric-pair regression (SPR) to estimate the asymptotic variance from the same private messages used for point estimation, without holding out participants or requesting a second release. We prove the central limit theorem and consistency of the fully transcript-based sandwich covariance and intrinsic Wald region under local differential privacy. Simulations across various statistical problems and manifolds support the predicted decrease in estimation error and near-nominal coverage under moderate privacy. An application to NHANES anthropometric data illustrates private estimation of a leading body-size direction and its uncertainty.

stat.ML↗

Asymptotic Anytime-Valid Quantile Inference under Local Differential Privacy

Sequential quantile inference is difficult under local differential privacy because every record is randomized before reaching the analyst and the limiting quantile variance depends on an unknown density. We develop an online procedure that combines randomized response with dynamically chained parallel stochastic gradient descent (P-SGD). The resulting Polyak--Ruppert estimator admits a strong Gaussian approximation. A cross-chain quadratic statistic, computed entirely from private iterates, consistently estimates the limiting variance without a separate online density estimator. These results yield asymptotic confidence sequences and, under polynomial chain growth, asymptotic time-uniform coverage. Arm-wise constructions support locally private quantile best-arm identification, time-uniform simple-regret bounds, and sequential A/B tests of quantile treatment effects. Simulations and salary-data analyses illustrate the finite-sample behavior and practical use of the proposed methods.

stat.ME↗

FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence

Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning methods are rapidly expanding the design space of embodied policies, yet turning these algorithms into reliable robot systems remains constrained by fragmented data formats, training stacks, evaluation protocols, inference runtimes, and embodiment-specific interfaces. We present $\mathrm{FluxVLA}$ Engine, an open, configuration-driven platform that turns heterogeneous embodied-policy components into a reproducible data-to-deployment workflow. Rather than introducing another policy model, $\mathrm{FluxVLA}$ standardizes interfaces for datasets, visual-language and world models, action heads, reward- or advantage-weighted learning, distributed training, simulation evaluation, optimized inference, and robot operators. The engine further integrates compositional dual-arm simulation, scalable automatic data generation, and model-decoupled human-in-the-loop rollout, takeover, correction collection, and reward annotation. For responsive physical execution, it combines Real-Time Chunking (RTC) with accelerated inference backends, lightweight remote GPU serving, and configurable trajectory post-processing. Together, these capabilities connect offline learning, simulation validation, online correction, and real-robot execution through shared and auditable contracts. $\mathrm{FluxVLA}$ therefore targets the engineering bottlenecks separating promising embodied-learning algorithms from reproducible evaluation and dependable deployment. Code is available at https://github.com/FluxVLA/FluxVLA

cs.RO↗

Multicollinearity-agnostic feature screening for non-Euclidean responses: a factor adjusted approach

In high-dimensional settings, multicollinearity is a pervasive issue that can substantially impair the performance of feature screening methods based on marginal Fréchet regression. Feature screening for non-Euclidean responses becomes unreliable when ultrahigh-dimensional predictors suffer from multicollinearity, because feature-specific signals may be masked by shared latent factors. To mitigate this effect, we propose a Factor adjusted Fréchet sure independence screening procedure. The method first recovers latent common factors from the predictors and then evaluates each feature by the incremental Fréchet coefficient of determination contributed by its idiosyncratic component beyond the common factors. Under regularity conditions, we establish uniform approximation rates for the feasible screening utilities and prove the sure screening and sure ranking properties. Extensive numerical experiments provide compelling empirical support for the validity and effectiveness of our approach, particularly in scenarios with highly correlated covariates. We further illustrate the practical performance of our method through two representative non-Euclidean datasets: the ADNI dataset and the mortality dataset, both with distribution-valued responses.

stat.ME↗

MATCH: Multiplier-Assisted Tests for Conditional Hypotheses in Non-Euclidean Data

We propose a new procedure MATCH (Multiplier-Assisted Tests for Conditional Hypotheses) to test whether the non-Euclidean data match the target model, which is a general framework for significance and specification testing in Fréchet regression. MATCH covers global significance, partial significance, and the adequacy of global Fréchet regression, providing a unified way to compare unrestricted conditional Fréchet means with restricted alternatives. One of the key challenges is that the ordinary held-out loss difference is first-order degenerate under the null: the oracle losses coincide, and plug-in statistics is dominated by nuisance estimation error. MATCH uses sample splitting and independent random multipliers on held-out losses to create a nondegenerate Gaussian leading term without residuals or tangent-space coordinates. To improve data use and stability, we further develop cross-fitted tests and repeated cross-fitting with p-value merging. We establish asymptotic null validity, consistency under fixed alternatives, and local power guarantees. Simulations for distributional, symmetric positive-definite (SPD) matrix-valued, and spherical responses support the theoretical findings, and applications to county-level household income distributions and North Atlantic tropical-cyclone locations demonstrate the practical use of the proposed tests.

stat.ME↗

Locally Private Online Quantile Regression: Estimation and Inference

We study estimation and inference for online quantile regression under a one-report user-level $\eps$-locally differentially private ($\eps$-LDP) protocol. The main difficulty is that the standard quantile-regression estimating-equation contribution couples covariates with a residual comparison, so a server that receives only privatized reports cannot form the usual online update. We address this by developing a finite-alphabet channel in which each user computes the contribution locally, applies support-aware stochastic quantization and randomized response to one selected-block category, and sends one report. A public decoder corrects the randomized-response distortion and reconstructs a server-side estimating-equation input with the correct conditional mean. These decoded inputs are then used in projected Polyak-Ruppert averaging. For fixed finite channel designs, we establish local privacy, decoder unbiasedness, consistency, asymptotic normality, and Hessian-free self-normalized inference for prespecified scalar contrasts. Simulations and a New York City taxi-trip illustration show that the private trajectory approaches the nonprivate online reference as the privacy budget grows and outperforms direct Laplace and face-exponential geometric releases in the reported regimes.

stat.ML↗

Differentially private inference framework for Riemannian manifold data

We propose a novel and systematic differentially private (DP) inference framework for non-Euclidean data. First, we design two types of DP mechanisms for the Fréchet mean and variance for i.i.d. Riemannian manifold-valued data, tailored to different geometric structures and accompanied by analytic privacy budgets calibrated to the geometry of the underlying manifold. Second, we establish the consistency and central limit theorems (CLTs) of the proposed DP estimators, enabling a suite of statistical inference procedures under privacy constraints. Furthermore, we provide comprehensive implementation guidelines and feasible procedures, including consistent DP estimators of the asymptotic variance in the CLTs. Extensive numerical experiments support the proposed methodologies. Finally, we demonstrate the effectiveness of our approach on real-world medical image and sociological datasets supported on two representative manifolds.

stat.ME↗

Unified theory of testing relevant hypotheses in functional time series

In this paper, we develop a {\em unified} framework for testing relevant hypotheses in functional time series. The proposed approach accommodates one-sample, two-sample, and change point problems for contaminated observations under arbitrary sampling schemes. Combining B-spline estimation with self-normalization, we construct nuisance-parameter-free tests that bypass auxiliary estimation of long-run covariance functions and measurement-error variance functions. We establish asymptotic validity by exploiting a sequential Gaussian approximation for dependent random vectors of moderately high dimension, which leads to a pivotal limiting distribution. We also provide sufficient conditions for the non-degeneracy of the self-normalizer and establish consistent decision rules. A key theoretical finding is that the proposed tests detect \(n^{-1/2}\)-local alternatives under arbitrary sampling frequencies. This uncovers a sparse-to-dense phase transition distinct from those typically observed in functional data analysis: while the sampling frequency affects the asymptotic variance, the detection rate remains \(n^{-1/2}\), even in sparsely sampled regimes. We further study multiple change point alternatives and extend the theory to settings where consistent change point estimates are available. We also discuss the choice of self-normalizers, including the recently developed range-adjusted self-normalizer. Extensive simulations support the theoretical results, and applications to the AU.SHF implied volatility and traffic volume datasets demonstrate the practical utility of the proposed methods.

stat.ME↗

Asymptotic Anytime-Valid Inference for U-statistics

We study asymptotic anytime-valid confidence sequences for degree-two U-statistics under continuous monitoring. In the nondegenerate case, Hoeffding's projection reduces the problem to a time-uniform central limit theory for the partial sums of the first-order projection, while the canonical remainder is shown to be negligible under mild moment assumptions. A leave-one-out jackknife estimator then yields a fully data-driven procedure, leading to confidence sequences with asymptotic coverage guarantee for the parameter of interest. In the degenerate case, we show that the U-statistic is approximated by a centered quadratic Gaussian-chaos rather than by a simple Gaussian, which poses significant challenges for sequential inference. To address this issue, we novelly develop the Spectrally Allocated Gaussian-chaos Excursion (SAGE) boundary, and then provide plug-in implementations based on truncated spectrum estimation with consistency guarantees. The resulting widths can attain the expected time-uniform optimal rates: $\sqrt{\log\log n/n}$ in the nondegenerate regime and $\log\log n/n$ in the degenerate regime. Several widely used U-statistics are discussed within the proposed framework, and numerical experiments further support the validity of the derived theory.

math.ST↗

Geometric Renyi Differential Privacy: Ricci Curvature Characterized by Heat Diffusion Mechanisms

In this paper, we develop a novel privacy mechanism for Riemannian manifold-valued data. Our key contribution lies in uncovering unexpected connections among geometric analysis, heat diffusion models, and differential privacy (DP). We characterize the Renyi divergence via dimension-free Harnack inequalities on Riemannian manifolds and establish Renyi differential privacy guarantees governed by Ricci curvature. For manifolds with nonnegative Ricci curvature, we propose a mechanism based on heat diffusion. In contrast, for general manifolds we introduce a Langevin-process-based approach that yields intrinsic mechanisms supporting normalization-free sampling and continuous privacy-utility trade-offs. We derive detailed utility analyses for both mechanisms. As a statistical application, we develop privacy-preserving estimation of the generalized Frechet mean, including nontrivial sensitivity analysis and phase transition characterizations. Numerical experiments further demonstrate the advantages of the proposed DP mechanisms over existing approaches.

stat.ML↗

ARM: Advantage Reward Modeling for Long-Horizon Manipulation

Long-horizon robotic manipulation remains challenging for reinforcement learning (RL) because sparse rewards provide limited guidance for credit assignment. Practical policy improvement thus relies on richer intermediate supervision, such as dense progress rewards, which are costly to obtain and ill-suited to non-monotonic behaviors such as backtracking and recovery. To address this, we propose Advantage Reward Modeling (ARM), a framework that shifts from hard-to-quantify absolute progress to estimating relative advantage. We introduce a cost-effective tri-state labeling strategy -- Progressive, Regressive, and Stagnant -- that reduces human cognitive overhead while ensuring high cross-annotator consistency. By training on these intuitive signals, ARM enables automated progress annotation for both complete demonstrations and fragmented DAgger-style data. Integrating ARM into an offline RL pipeline allows for adaptive action-reward reweighting, effectively filtering suboptimal samples. Our approach achieves a 99.4% success rate on a challenging long-horizon towel-folding task, demonstrating improved stability and data efficiency over current VLA baselines with near-zero human intervention during policy training.

cs.RO↗

Strong Gaussian approximation for U-statistics in high dimensions and beyond

We establish a strong Gaussian approximation for high-dimensional non-degenerate U-statistics with diverging dimension. Under mild assumptions, we construct, on a sufficiently rich probability space, a Gaussian process that uniformly approximates the entire sequential U-statistic process. The approximation error is explicitly characterized and vanishes under polynomial growth of the dimension. The key technical contribution is a sharp martingale maximal inequality for completely degenerate U-statistics, combined with a high-dimensional strong approximation for independent sums. This coupling yields functional Gaussian limits without relying on $\mathcal{L}^\infty$-type bounds or bootstrap arguments. The theory is illustrated through three representative examples of U-statistics: the spatial Kendall's tau matrix, the multivariate Gini's mean difference, and the characteristic dispersion parameter. As applications, we derive Brownian bridge approximations for U-statistic-based change-point statistics and develop a self-normalized relevant testing procedure whose limiting distribution is fully pivotal. The framework naturally accommodates bounded kernels and therefore remains valid under heavy-tailed distributions. Overall, our results provide a unified probability-theoretic foundation for high-dimensional inference based on U-statistics.

math.ST↗

Individualized Causal Effects under Network Interference with Combinatorial Treatments

Modern causal decision-making increasingly demands individualized treatment-effect estimation in networks where interventions are high-dimensional, combinatorial vectors. While network interference, effect heterogeneity, and multi-dimensional treatments have been studied separately, their intersection yields an exponentially large intervention space that makes standard identification tools and low-dimensional exposure mappings untenable. We bridge this gap with a unified framework that constructs a \emph{global potential-outcome emulator} for unit-level inference. Our method combines (1) rooted network configurations to leverage local smoothness, (2) doubly robust orthogonalization to mitigate confounding from network position and covariates, and (3) sparse spectral learning to efficiently estimate response surfaces over the $2^p$-dimensional treatment space. We also decompose networked effects into own-treatment, structural, and interaction components, and provide finite-sample error bounds and asymptotic consistency guarantees. Overall, we show that individualized causal inference remains feasible in high-dimensional networked settings without collapsing the intervention space.

stat.ME↗

LightCity: An Urban Dataset for Outdoor Inverse Rendering and Reconstruction under Multi-illumination Conditions

Inverse rendering in urban scenes is pivotal for applications like autonomous driving and digital twins. Yet, it faces significant challenges due to complex illumination conditions, including multi-illumination and indirect light and shadow effects. However, the effects of these challenges on intrinsic decomposition and 3D reconstruction have not been explored due to the lack of appropriate datasets. In this paper, we present LightCity, a novel high-quality synthetic urban dataset featuring diverse illumination conditions with realistic indirect light and shadow effects. LightCity encompasses over 300 sky maps with highly controllable illumination, varying scales with street-level and aerial perspectives over 50K images, and rich properties such as depth, normal, material components, light and indirect light, etc. Besides, we leverage LightCity to benchmark three fundamental tasks in the urban environments and conduct a comprehensive analysis of these benchmarks, laying a robust foundation for advancing related research.

cs.CV↗

Federated Learning of Quantile Inference under Local Differential Privacy

In this paper, we investigate federated learning for quantile inference under local differential privacy (LDP). We propose an estimator based on local stochastic gradient descent (SGD), whose local gradients are perturbed via a randomized mechanism with global parameters, making the procedure tolerant of communication and storage constraints without compromising statistical efficiency. Although the quantile loss and its corresponding gradient do not satisfy standard smoothness conditions typically assumed in existing literature, we establish asymptotic normality for our estimator as well as a functional central limit theorem. The proposed method accommodates data heterogeneity and allows each server to operate with an individual privacy budget. Furthermore, we construct confidence intervals for the target value through a self-normalization approach, thereby circumventing the need to estimate additional nuisance parameters. Extensive numerical experiments and real data application validate the theoretical guarantees of the proposed methodology.

stat.ME↗

Panoramic Direct LiDAR-assisted Visual Odometry

Enhancing visual odometry by exploiting sparse depth measurements from LiDAR is a promising solution for improving tracking accuracy of an odometry. Most existing works utilize a monocular pinhole camera, yet could suffer from poor robustness due to less available information from limited field-of-view (FOV). This paper proposes a panoramic direct LiDAR-assisted visual odometry, which fully associates the 360-degree FOV LiDAR points with the 360-degree FOV panoramic image datas. 360-degree FOV panoramic images can provide more available information, which can compensate inaccurate pose estimation caused by insufficient texture or motion blur from a single view. In addition to constraints between a specific view at different times, constraints can also be built between different views at the same moment. Experimental results on public datasets demonstrate the benefit of large FOV of our panoramic direct LiDAR-assisted visual odometry to state-of-the-art approaches.

cs.RO↗