Search arXiv⌕ Search

arXiv subjects

Aaditya Ramdas

Publications and source records attributed to Aaditya Ramdas.

At least 19 recordsLinked to original sources

Distribution-uniform strong laws of large numbers

We revisit the question of whether the strong law of large numbers (SLLN) holds uniformly in a rich family of distributions, culminating in a distribution-uniform generalization of the Marcinkiewicz-Zygmund SLLN. These results can be viewed as extensions of Chung's distribution-uniform SLLN to random variables with uniformly integrable $q^\text{th}$ absolute central moments for $0 < q < 2$. Furthermore, we show that uniform integrability of the $q^\text{th}$ moment is both sufficient and necessary for the SLLN to hold uniformly at the Marcinkiewicz-Zygmund rate of $n^{1/q - 1}$. These proofs centrally rely on novel distribution-uniform analogues of some familiar almost sure convergence results including the Khintchine-Kolmogorov convergence theorem, Kolmogorov's three-series theorem, a stochastic generalization of Kronecker's lemma, and the Borel-Cantelli lemmas. We also consider the non-identically distributed case.

math.PR↗

Nonasymptotic and distribution-uniform Komlós-Major-Tusnády approximation

We present nonasymptotic concentration inequalities for sums of independent and identically distributed random variables that yield asymptotic strong Gaussian approximations of Komlós, Major, and Tusnády (KMT) [1975,1976]. The constants appearing in our inequalities are either universal or explicit, and thus as corollaries, they imply distribution-uniform generalizations of the aforementioned KMT approximations. In particular, it is shown that uniform integrability of a random variable's $q^{\text{th}}$ moment is both necessary and sufficient for the KMT approximations to hold uniformly at the rate of $o(n^{1/q})$ for $q > 2$ and that having a uniformly lower bounded Sakhanenko parameter -- equivalently, a uniformly upper-bounded Bernstein parameter -- is both necessary and sufficient for the KMT approximations to hold uniformly at the rate of $O(\log n)$. Instantiating these uniform results for a single probability space yields the analogous results of KMT exactly.

math.PR↗

Conformal Prediction Through the Lens of Hypothesis Testing: Universality, Impossibility, and Optimality

The connections between conformal prediction and permutation tests are already widely-known in the literature. Some authors motivate conformal prediction by saying that it computes a permutation p-value for the hypothesis $H_0 : Y_{n+1} = y$, and then inverts this to form a prediction set for $Y_{n+1}$ (i.e., accepts all values $y$ into the prediction set for which the p-value is large). In this paper, we examine an alternative view, which is less well-known: we again cast conformal prediction via the inversion of a permutation test, but for the null of exchangeability of the joint distribution of the $n+1$ samples. This change in perspective, while simple, adheres more closely to traditional formalization in hypothesis testing, which offers several benefits. First, we use the duality between conformal sets and testing to show that foundational universality and impossibility results in the conformal prediction literature can be reproduced directly using classical hypothesis testing theory (due to Neyman, Lehmann, Scheff{é}, Kraft, Le Cam, and others). Furthermore, we show that an optimality result for conformal prediction can be derived using standard Neyman-Pearson theory: for any joint distribution of the covariates and response $X,Y$, and any sample size, the optimal method for prediction sets---delivering the most efficient set among all methods with valid coverage for exchangeable distributions---is a conformal predictor whose score is the inverse conditional density of $Y|X$.

math.ST↗

Change detection with conformal martingales: new optimal constructions, and suboptimality of existing methods

We study distribution-free sequential changepoint detection for independent observations with unknown and unrestricted pre- and post-change laws. We build on the conformal test martingales and associated e-detectors of Vovk(2021), which control the probability of false alarm (PFA) and the average run length (ARL) respectively. The majority of these works focus on validity, with statistical efficiency usually left for simulations. We develop a comprehensive theory of how conformal p-values behave under non-exchangeable data with a changepoint at an unknown time $T$. We use this to analyze the post-change growth and resulting detection delay of conformal martingale methods, and prove that the standard existing methods are suboptimal for PFA and ARL control, and can lead to delays that are $Ω(T)$ and $Ω(\sqrt{\text{ARL}})$ respectively. We propose different conformal e-processes and e-detectors that are provably minimax optimal, with delays $Θ(\log T)$ and $Θ(\log \text{ARL})$ respectively, and have much shorter delays in simulations.

math.ST↗

A variational approach to dimension-free self-normalized concentration

We study the self-normalized concentration of vector-valued stochastic processes. We focus on bounds for "sub-$ψ$" processes, a well-known and quite general class that encompasses a wide variety of well-known tail conditions (including sub-exponential, sub-Gaussian, sub-gamma, sub-Poisson, and several heavy-tailed settings without a moment generating function such as symmetric or bounded 2nd or 3rd moments). Our results recover and generalize the influential bound of de la Peña et al. [20] (proved again in Abbasi-Yadkori et al. [2]) in the sub-Gaussian case. Further, we fill a gap in the literature between determinant-based bounds and more recent bounds based on condition numbers. As applications we prove a Bernstein inequality for random vectors satisfying a moment condition (a more general condition than boundedness), and also provide the first dimension-free self-normalized empirical Bernstein inequality. Our techniques are based on the variational (PAC-Bayes) approach to concentration.

math.PR↗

Extreme classification: beating chance with one training example from each class

We study a minimal classification problem: Given independent labeled observations $X\sim P$ and $Z\sim Q$ from two unknown distributions $P,Q$, and given an independent target $Y$ drawn with equal probability from $P$ or $Q$, can one classify $Y$ strictly better than chance whenever $P\neq Q$? The one-nearest-neighbor rule succeeds for every pair of multivariate Gaussian distributions with distinct means and a common positive-definite covariance matrix but can perform strictly worse than chance even for smooth densities on the real line. We construct a fixed randomized kernel rule whose expected accuracy is exactly $1/2+\operatorname{MMD}_k^2(P,Q)/4$, and obtain characteristic kernels on countably generated measurable spaces from countable families of measurable binary questions. We also prove that a deterministic order rule on $\mathbb R$ beats chance for every pair of distinct Borel probability measures. A measurable encoding then gives a deterministic distribution-free rule which beats chance on every countably generated measurable space, in particular every separable metric space. Finally, we show that no rule works for every distinct pair of distributions and every unknown unbalanced class prior; under adaptive target-class selection, every rule other than a fair coin is strictly worse than chance for some finitely supported pair.

stat.ML↗

Distribution-free inference on the number of changepoints

Suppose we are given an ordered sequence of independent data whose distribution changes $K$ times at unknown locations, for some unknown $K \geq 0$. In this paper, we study the problem of performing distribution-free inference on $K$. First, we show an impossibility result: any distribution-free upper confidence bound on $K$ must be trivial and uninformative. Then, using conformal $p$-values, and under only the assumption that the data segments induced by the changepoints are exchangeable (within themselves) and mutually independent, we construct a finite-sample valid lower confidence bound on $K$, which we call the Conformal LOwer bound on Changepoint Count (CLOCC). We show that CLOCC is the only feasible way to provide a lower bound on $K$ under the stated assumptions, a property we refer to as its universality. We provide practical guidelines for choosing score functions that yield efficient and tight lower bounds. We evaluate CLOCC in several synthetic and real-data experiments, where it provides informative lower bounds on $K$, demonstrating its practical applicability.

stat.ML↗

A complete characterization of sequential testability and change detectability in i.i.d. models

We give a necessary and sufficient condition for the existence of power-one sequential tests in an i.i.d. composite testing problem. A level-\(α\) test with power one against every alternative exists if and only if the alternatives are separated from the null by a countable family of finite-block events. We provide other equivalent conditions using randomized fixed-sample tests, bounded finite-block scores, e-processes, reduced-filtration test supermartingales, and a countable cover whose finite-block weak-$*$ closed convex hulls are positively separated in total variation. As a bonus, the constructive proof yields tests have pointwise expected sample size \(O_Q(\log(1/α))\). Exactly the same conditions also characterize i.i.d.\ change detectability under optional-horizon average-run-length control: for every \(η>0\), they are equivalent to an alarm family \((T_γ)_{γ\ge1}\) satisfying \(\Prob_{P^\infty}(T_γ\leσ)\le \E_{P^\infty}σ/γ\) for every null law and every stopping time \(σ\). In fact, when these conditions hold, we can construct a single e-detector such that every null-law average run length lies between \(γ\) and \((1+η)γ+1\), and having robust Lorden delay \(O_Q(\logγ)\).

math.ST↗

A composite generalization of Ville's martingale theorem using e-processes

We provide a composite version of Ville's theorem that an event has zero measure if and only if there exists a nonnegative martingale which explodes to infinity when that event occurs. This is a classic result connecting measure-theoretic probability to the sequence-by-sequence game-theoretic probability, recently developed by Shafer and Vovk. Our extension of Ville's result involves appropriate composite generalizations of nonnegative martingales and measure-zero events: these are respectively provided by ``e-processes'', and a new inverse capital outer measure. We then develop a novel line-crossing inequality for sums of random variables which are only required to have a finite first moment, which we use to prove a composite version of the strong law of large numbers (SLLN). This allows us to show that violation of the SLLN is an event of outer measure zero and that our e-process explodes to infinity on every such violating sequence, while this is provably not achievable with a nonnegative (super)martingale.

math.PR↗

E-values as statistical evidence: A comparison to Bayes factors, likelihoods, and p-values

A recurring debate in the philosophy of statistics concerns what, exactly, should count as a measure of evidence for or against a given hypothesis. P-values, likelihood ratios, and Bayes factors all have their defenders. In this paper we add two additional candidates to this list: the e-value and its sequential analogue, the e-process. E-values enjoy several desirable properties as measures of evidence: they combine naturally across studies, handle composite hypotheses, provide long-run error rates, and admit a useful interpretation as the wealth accrued by a bettor in a game against the null distribution. E-processes additionally handle optional stopping and optional continuation. This work examines the extent to which e-values and e-processes satisfy the evidential desiderata of different statistical traditions, concluding that they combine attractive features of p-values, likelihood ratios, and Bayes factors, and merit serious consideration as interpretable and intuitive measures of statistical evidence.

stat.ME↗

Power one sequential tests exist for weakly compact $\mathscr P$ against $\mathscr P^c$

We study power-one sequential testing for an i.i.d. law on a Polish sample space. Given a nonempty composite null class $\Pcal\subseteq\mathcal M_1(\X)$, we ask when there exists a level-$α$ stopping rule that rejects almost surely under every alternative in a prescribed class $\Qcal\subseteq\Pcal^c$. Our main sufficient condition is local weak lower semicontinuity and positivity of the information projection functional \( Φ_\Pcal(Q):=\inf_{P\in\Pcal}\KL(Q\|P). \) In particular, if $\Pcal$ is weakly compact, then for every $α\in(0,1)$ there is a single level-$α$ sequential test with power one against the entire complement $\Pcal^c$. The proof combines Csiszár's nonasymptotic Sanov bound for weakly closed convex empirical-measure sets with a Lindelöf countable-subcover argument. We also show that weak lower semicontinuity is sufficient but not necessary by giving examples where discontinuous finite-sample events separate alternatives that weak neighborhoods cannot detect. Finally, we construct an $e$-process that is asymptotically relatively growth-rate optimal under weak compactness. We verify the weak-lower-semicontinuity condition for weakly compact nulls, $f$-divergence balls, several integral probability metric balls, Wasserstein balls on proper spaces, and a number of non-weakly-compact semiparametric examples.

math.ST↗

Gaussian-efficient testing by betting on the mean of bounded data

Given $[0,1]$-valued random variables $X_1,\dots,X_n$ such that $\mathbb{E}[X_i | X_1,\dots,X_{i-1}]= μ$ for all $i$, we propose a new nonasymptotic confidence interval for $μ$ that is obtained by inverting terminal e-values generated by a novel betting strategy. When the data are iid, its limiting width matches that of the central limit theorem (``Gaussian-efficient''), finally surpassing the inefficient limits of previous betting intervals. Our main conceptual advance involves designing betting fractions that track the conditional rejection probability of the most powerful terminal test in a limiting Gaussian experiment. When one predictable variance estimator is shared across candidate means, the deterministic inversion is an interval for every data sequence and its two endpoints can be found easily. The width can be improved further with external randomization. In simulations, our method yields the tightest intervals to date; for every distribution tested and all sufficiently large $n$, our deterministic version beats STaR-Bets and is competitive with Gaffke, while the randomized improvement beats both. It thus combines finite-sample validity under martingale dependence, easy endpoint computation, Gaussian-efficient inference for iid data, and excellent empirical performance. We also extend the construction and its efficiency theory to sampling without replacement, where it again achieves state-of-the-art empirical performance.

stat.ME↗

On Stopping Times of Power-one Sequential Tests: Tight Lower and Upper Bounds

We present two general lower bounds for stopping times of sequential tests between arbitrary composite nulls $\mathcal P$ and alternatives $\mathcal Q$. The first lower bound is for the ``Wald setting'' where the type-1 error level $α$ approaches zero for a fixed alternative $Q \in \mathcal Q$, and equals $\log(1/α)$ divided by a certain infimum KL divergence between $\mathcal P$ and $Q$, termed $\operatorname{KL_{inf}}$. The second lower bound applies to the ``Farrell setting'', where $α$ is fixed and $\operatorname{KL_{inf}}$ approaches $0$ along a sequence of alternatives such that the required expected sample size along that sequence is of order at least $\operatorname{KL^{-1}_{inf}} \log \log \operatorname{KL^{-1}_{inf}}$. Our main contribution is the generality of these bounds, which hold in non-parametric, composite settings, without requiring a dominating reference measure, substantially generalizing the known parametric results. We also provide sufficient conditions for matching upper bounds and show that these are met in several nontrivial non-parametric cases.

math.ST↗

Universality of e-detectors for ARL control

An e-detector for a pre-change class $\mathcal P$ is a nonnegative process $M$ such that $\mathbb E_P[M_τ] \leq \mathbb E_P[τ]$ for all stopping times $τ$ and all $P \in \mathcal P$. Thresholding e-detectors controls the average run length (ARL): declaring a change at the first time $T_b$ when $M$ crosses $b$ ensures that $\inf_{P \in \mathcal P}\mathbb E_P[T] \geq b$. But e-detectors do substantially more than control the ARL; they also satisfy a \emph{optional-horizon inequality}: \[ P(T_b\leqσ)\leq \mathbb E_P[σ]/b \] for every data-dependent stopping time (monitoring horizon) \(σ\) and $P\in \mathcal P$. In particular, every e-detector-based procedure obeys $P(T\leq t)\leq t/b$ at each fixed $t$, thus avoiding early false alarms. Remarkably, the converse also holds: every stopping time $T$ that satisfies the optional-horizon inequality must in fact arise from thresholding an e-detector. We also derive a universal representation of stopping times that satisfy (only) ARL control. These are represented by \emph{weak} e-detectors, that only require $\mathbb E_P[M_τ] \leq \mathbb E_P[τ]$ to hold at all threshold stopping times $T_b$. Appendices present universal representations for other (less common) change detection metrics.

math.ST↗

The dynamic generalized covariance measure for conditional independence testing with nonstationary time series

Identifying relationships among stochastic processes is a core objective in many fields, such as economics. While the standard toolkit for multivariate time series analysis has many advantages, it can be difficult to capture nonlinear dynamics using linear vector autoregressive models. This difficulty has motivated the development of methods for causal discovery and variable selection for nonlinear time series, which routinely employ tests for conditional independence. In this paper, we introduce the first framework for conditional independence testing that works with a single realization of a nonstationary nonlinear process. The proposed test is designed to have power against alternatives in which the expected conditional covariance is non-zero for at least some times. We also discuss an approach for gaining power against a broader range of alternatives. The key technical ingredients of our framework are time-varying nonlinear regression, estimation of local long-run covariance matrices of products of error processes, and a distribution-uniform strong Gaussian approximation.

stat.ME↗

Combining e-values using demi-supermartingales

We present a new method for combining e-variables through demi-supermartingales, which settles an old conjecture in the literature on nonparametric mean testing. It also provides an explicit concentration bound for a certain Kullback--Leibler-type statistic arising in the stochastic multi-armed bandit literature. All of these combination results hold for independent e-variables as well as for the class of co-valid e-variables, whose dependence structure lies somewhere between independence and sequential validity. The results are further generalized to compound e-variables. The proofs proceed by analyzing elementary symmetric polynomials and their behavior as nonnegative demi-supermartingales.

stat.ME↗

Bringing Closure to False Discovery Rate Control: A General Principle for Multiple Testing

We present a novel necessary and sufficient principle for multiple testing methods controlling an expected loss. This principle asserts that every such multiple testing method is a special case of a general closed testing procedure based on e-values. It generalizes the Closure Principle, known to underlie all methods controlling familywise error and tail probabilities of false discovery proportions, to a large class of error rates -- in particular to the false discovery rate (FDR). By writing existing methods as special cases of this procedure, we can achieve uniform improvements, as we demonstrate for the e-Benjamini-Hochberg and the Benjamini-Yekutieli procedures, and the self-consistent method of Su (2018). We also show that methods derived using our novel e-Closure Principle generally control their error rate not just for one rejected set, but simultaneously over many, allowing post hoc flexibility for the researcher. Moreover, because all multiple testing methods for the expected loss error metrics covered by our framework are derived from the same procedure, researchers may even choose the error metric post hoc. Under certain conditions, this flexibility even extends to post hoc choice of the nominal error rate.

stat.ME↗

Non-partitioned e-detectors for nonparametric sequential change detection

We study the problem of sequential change detection over a general class of probability distributions ($\mathcal P$), where both the pre-change and post-change distributions are unknown and belong to $\mathcal P$. We do not assume a pre-specified partition of $\mathcal P$ into pre- and post-change families. We propose a general class of sequential change detectors obtained by aggregating point-null e-processes over possible changepoints and taking an infimum over candidate no-change distributions. The weights in the aggregation scheme determine whether they attain average run length (ARL) control and probability-of-false-alarm (PFA) control. Under suitable assumptions, we prove that our methods achieve first-order asymptotically optimal detection delay. Concrete examples include sub-Gaussian and bounded mean changes, Gaussian mean changes with unknown variance, as well as changes in Markov transition matrices.

stat.ME↗