Search arXivSearch

arXiv subjects

Axel Munk

Publications and source records attributed to Axel Munk.

At least 19 recordsLinked to original sources

Gromov-Wasserstein Barycenter Surrogates: Statistical Methodology, Distributional Limits and Applications

We introduce statistical theory for the matching of finitely many objects, represented as metric measure spaces (mm-spaces). The approach is based on the second lower bound (SLB) of the Gromov-Wasserstein distance and thus is able to identify deviations in the distributions of the (pairwise) distances within each mm-space. We introduce a surrogate of the SLB barycenter which can be easily computed and expressed explicitly in terms of the distance distributions of each object. When comparing $m$ mm-spaces for $n$ randomly drawn samples in each space, the resulting statistic then can be calculated efficiently in $O(m \cdot n^2 \log(n))$ basic operations. We derive the asymptotic distribution and finite-sample bounds of the proposed test statistic, which serves as a basis for a variety of tools for statistical inference, specifically an asymptotic test for pose-invariant object discrimination and a classification method (based on the SLB barycenter) with controlled error rates. These methods are investigated in simulations and applied to the structural comparison of protein domains.

math.ST

Distributional Convergence of Empirical Entropic Optimal Transport and Statistical Applications

Recently, the statistical properties of empirical Entropic Optimal Transport (EOT) have attracted great interest, as this quantity has been shown to be useful for complex data analysis, among other reasons due to its computational efficiency. In several applications, it has been observed that the EOT plan provides valuable information beyond just the optimal value. For example, in cell biology, colocalization analysis based on the EOT plan has been introduced as a measure for quantification of spatial proximity of different protein assemblies. Despite recent progress in the analysis of its risk properties, a precise understanding of its statistical fluctuations to make it accessible for inference remains elusive to a large extent. In this paper, we derive asymptotic weak convergence result for a large class of functionals of the EOT plan, in which the colocalization process is included. The proof is based on Hadamard differentiability and the extended delta method. As an application, we obtain uniform confidence bands for colocalization curves and bootstrap consistency. Our theory is supported by simulation studies and is illustrated by real world data analysis from mitochondrial protein colocalization.

math.ST

Quantile characterization of univariate unimodality

Unimodal univariate distributions can be characterized as piecewise convex-concave cumulative distribution functions. In this note we transfer this shape constraint characterization to the quantile function. We show that this characterization comes with the upside that the quantile function of a unimodal distribution is always absolutely continuous and consequently unimodality is equivalent to the quasi-convexity of its Radon-Nikodym derivative, i.e., the quantile density. Our analysis is based on the theory of generalized inverses of non-decreasing functions and relies on a version of the inverse function rule for non-decreasing functions.

math.ST

Weak convergence of Bayes estimators under general loss functions

We investigate the asymptotic behavior of parametric Bayes estimators under a broad class of loss functions that extend beyond the classical translation-invariant setting. To this end, we develop a unified theoretical framework for loss functions exhibiting locally polynomial structure. This general theory encompasses important examples such as the squared Wasserstein distance, the Sinkhorn divergence and Stein discrepancies, which have gained prominence in modern statistical inference and machine learning. Building on the classical Bernstein--von Mises theorem, we establish sufficient conditions under which Bayes estimators inherit the posterior's asymptotic normality. As a by-product, we also derive conditions for the differentiability of Wasserstein-induced loss functions and provide new consistency results for Bayes estimators. Several examples and numerical experiments demonstrate the relevance and accuracy of the proposed methodology.

math.ST

Optimal Transport Based Testing in Factorial Designs

We introduce a general framework for testing statistical hypotheses in factorial designs for probability measures supported on finite spaces. The suggested methodology is based on the pairwise comparison of measures using optimal transport (OT). The formulation of hypotheses is intuitive: It is a direct extension of those underlying the analysis of variance (ANOVA) and its nonparametric counterparts to test for linear relationships between (discrete) probability measures in factorial designs. To this end, means or cumulative distribution functions simply will be replaced by measures. We derive under the null hypotheses and under (local) alternatives the asymptotic distribution of the corresponding empirical OT test statistic, which is the optimal value of a linear program with random objective function. It turns out that this requires to extend existing techniques from probability measures to signed measures, and we show directional Hadamard differentiability and the validity of the functional delta method. We discuss computational issues, permutation and bootstrap tests, and back up our findings with simulations. We illustrate our methodology on datasets from cellular biophysics and from biometric identification.

math.ST

Sharp Convergence Rates of Empirical Unbalanced Optimal Transport for Spatio-Temporal Point Processes

We statistically analyze empirical plug-in estimators for unbalanced optimal transport (UOT) formalisms, focusing on the Kantorovich-Rubinstein distance, between general intensity measures based on observations from spatio-temporal point processes. Specifically, we model the observations by two weakly time-stationary point processes with spatial intensity measures $\mu$ and $\nu$ over the expanding window $(0,t]$ as $t$ increases to infinity, and establish sharp convergence rates of the empirical UOT in terms of the intrinsic dimensions of the measures. We assume a sub-quadratic temporal growth condition of the variance of the process, which allows for a wide range of temporal dependencies. As the growth approaches quadratic, the convergence rate becomes slower. This variance assumption is related to the time-reduced factorial covariance measure, and we exemplify its validity for various point processes, including the Poisson cluster, Hawkes, Neyman-Scott, and log-Gaussian Cox processes. Complementary to our upper bounds, we also derive matching lower bounds for various spatio-temporal point processes of interest and establish near minimax rate optimality of the empirical Kantorovich-Rubinstein distance.

math.ST

Local Poisson Deconvolution for Discrete Signals

We analyze the statistical problem of recovering an atomic signal, modeled as a discrete uniform distribution $\mu$, from a binned Poisson convolution model. This question is motivated, among others, by super-resolution laser microscopy applications, where precise estimation of $\mu$ provides insights into spatial formations of cellular protein assemblies. Our main results quantify the local minimax risk of estimating $\mu$ for a broad class of smooth convolution kernels. This local perspective enables us to sharply quantify optimal estimation rates as a function of the clustering structure of the underlying signal. Moreover, our results are expressed under a multiscale loss function, which reveals that different parts of the underlying signal can be recovered at different rates depending on their local geometry. Overall, these results paint an optimistic perspective on the Poisson deconvolution problem, showing that accurate recovery is achievable under a much broader class of signals than suggested by existing global minimax analyses. Beyond Poisson deconvolution, our results also allow us to establish the local minimax rate of parameter estimation in Gaussian mixture models with uniform weights. We apply our methods to experimental super-resolution microscopy data to identify the location and configuration of individual DNA origamis. In addition, we complement our findings with numerical experiments on runtime and statistical recovery that showcase the practical performance of our estimators and their trade-offs.

math.ST

Online jump and kink detection in segmented linear regression: Statistical optimality meets computational efficiency

We consider the problem of sequential (online) estimation of a single change point in a piecewise linear regression model under a Gaussian setup. We demonstrate that certain CUSUM-type statistics attain the minimax optimal rates for localizing the change point. Our minimax analysis unveils an interesting phase transition from a jump (discontinuity in function values) to a kink (a change in slope). Specifically, for a jump, the minimax rate is of order $\log (n) / n$ , whereas for a kink it scales as $(\log (n) / n)^{1/3}$, given that the sampling rate is of order $1/n$. We further introduce an online algorithm based on these detectors, which optimally identifies both a jump and a kink, and is able to distinguish between them. Notably, the algorithm operates with constant computational complexity and requires only constant memory per incoming sample. Finally, we evaluate the empirical performance of our method on both simulated and real-world data sets. An implementation is available in the R package FLOC on GitHub.

math.ST

Adaptive monotonicity testing in sublinear time

Modern large-scale data analysis increasingly faces the challenge of achieving computational efficiency as well as statistical accuracy, as classical statistically efficient methods often fall short in the first regard. In the context of testing monotonicity of a regression function, we propose FOMT (Fast and Optimal Monotonicity Test), a novel methodology tailored to meet these dual demands. FOMT employs a sparse collection of local tests, strategically generated at random, to detect violations of monotonicity scattered throughout the domain of the regression function. This sparsity enables significant computational efficiency, achieving sublinear runtime in most cases, and quasilinear runtime (i.e., linear up to a log factor) in the worst case. In contrast, existing statistically optimal tests typically require at least quadratic runtime. FOMT's statistical accuracy is achieved through the precise calibration of these local tests and their effective combination, ensuring both sensitivity to violations and control over false positives. More precisely, we show that FOMT separates the null and alternative hypotheses at minimax optimal rates over H\"older function classes of smoothness order in $(0,2]$. Further, when the smoothness is unknown, we introduce an adaptive version of FOMT, based on a modified Lepskii principle, which attains statistical optimality and meanwhile maintains the same computational complexity as if the intrinsic smoothness were known. Extensive simulations confirm the competitiveness and effectiveness of both FOMT and its adaptive variant.

math.ST

Identifiability and Exact Reconstruction of the Optimal Transport Cost on Finite Spaces

The goal of optimal transport (OT) is to find optimal assignments or matchings between data sets which minimize the total cost for a given cost function. However, sometimes the cost function is unknown but we have access to (parts of) the solution to the OT problem, e.g.\ the OT plan or the value of the objective function. Recovering the cost from such information is called inverse OT and has become recently of certain interest triggered by novel applications, e.g.\ in social science and economics. This raises the issue under which circumstances such cost is identifiable, i.e., it can be uniquely recovered from other OT quantities. In this work we provide sufficient and necessary conditions for the identifiability of the cost function on finite ground spaces. We find that such conditions correspond to the combinatorial structure of the corresponding linear program and discuss its computational complexity and implications for cost estimation in statistical linear models.

math.OC

Distributional limits of graph cuts on discretized grids

Graph cuts are among the most prominent tools for clustering and classification analysis. While intensively studied from geometric and algorithmic perspectives, graph cut-based statistical inference still remains elusive to a certain extent. Distributional limits are fundamental in understanding and designing such statistical procedures on randomly sampled data. We provide explicit limiting distributions for balanced graph cuts in general on a fixed but arbitrary discretization. In particular, we show that Minimum Cut, Ratio Cut and Normalized Cut behave asymptotically as the minimum of Gaussians as sample size increases. Interestingly, our results reveal a dichotomy for Cheeger Cut: The limiting distribution of the optimal objective value is the minimum of Gaussians only when the optimal partition yields two sets of unequal volumes, while otherwise the limiting distribution is the minimum of a random mixture of Gaussians. Further, we show the bootstrap consistency for all types of graph cuts by utilizing the directional differentiability of cut functionals. We validate these theoretical findings by Monte Carlo experiments, and examine differences between the cuts and the dependency on the underlying distribution. Additionally, we expand our theoretical findings to the Xist algorithm, a computational surrogate of graph cuts recently proposed in Suchan, Li and Munk (arXiv, 2023), thus demonstrating the practical applicability of our findings e.g. in statistical tests.

math.ST

Robust inference of cooperative behaviour of multiple ion channels in voltage-clamp recordings

Recent experimental studies have shed light on the intriguing possibility that ion channels exhibit cooperative behaviour. However, a comprehensive understanding of such cooperativity remains elusive, primarily due to limitations in measuring separately the response of each channel. Rather, only the superimposed channel response can be observed, challenging existing data analysis methods. To address this gap, we propose IDC (Idealisation, Discretisation, and Cooperativity inference), a robust statistical data analysis methodology that requires only voltage-clamp current recordings of an ensemble of ion channels. The framework of IDC enables us to integrate recent advancements in idealisation techniques and coupled Markov models. Further, in the cooperativity inference phase of IDC, we introduce a minimum distance estimator and establish its statistical guarantee in the form of asymptotic consistency. We demonstrate the effectiveness and robustness of IDC through extensive simulation studies. As an application, we investigate gramicidin D channels. Our findings reveal that these channels act independently, even at varying applied voltages during voltage-clamp experiments. An implementation of IDC is available from GitLab.

stat.ME

Nonlinear Inverse Optimal Transport: Identifiability of the Transport Cost from its Marginals and Optimal Values

The inverse optimal transport problem is to find the underlying cost function from the knowledge of optimal transport plans. While this amounts to solving a linear inverse problem, in this work we will be concerned with the nonlinear inverse problem to identify the cost function when only a set of marginals and its corresponding optimal values are given. We focus on absolutely continuous probability distributions with respect to the $d$-dimensional Lebesgue measure and classes of concave and convex cost functions. Our main result implies that the cost function is uniquely determined from the union of the ranges of the gradients of the optimal potentials. Since, in general, the optimal potentials may not be observed, we derive sufficient conditions for their identifiability - if an open set of marginals is observed, the optimal potentials are then identified via the value of the optimal costs. We conclude with a more in-depth study of this problem in the univariate case, where an explicit representation of the transport plan is available. Here, we link the notion of identifiability of the cost function with that of statistical completeness.

math.OC

A scalable clustering algorithm to approximate graph cuts

Due to their computational complexity, graph cuts for cluster detection and identification are used mostly in the form of convex relaxations. We propose to utilize the original graph cuts such as Ratio, Normalized or Cheeger Cut to detect clusters in weighted undirected graphs by restricting the graph cut minimization to $st$-MinCut partitions. Incorporating a vertex selection technique and restricting optimization to tightly connected clusters, we combine the efficient computability of $st$-MinCuts and the intrinsic properties of Gomory-Hu trees with the cut quality of the original graph cuts, leading to linear runtime in the number of vertices and quadratic in the number of edges. Already in simple scenarios, the resulting algorithm Xist is able to approximate graph cut values better empirically than spectral clustering or comparable algorithms, even for large network datasets. We showcase its applicability by segmenting images from cell biology and provide empirical studies of runtime and classification rate.

cs.DS

Multiscale scanning with nuisance parameters

We develop a multiscale scanning method to find anomalies in a $d$-dimensional random field in the presence of nuisance parameters. This covers the common situation that either the baseline-level or additional parameters such as the variance are unknown and have to be estimated from the data. We argue that state of the art approaches to determine asymptotically correct critical values for multiscale scanning statistics will in general fail when such parameters are naively replaced by plug-in estimators. Instead, we suggest to estimate the nuisance parameters on the largest scale and to use (only) smaller scales for multiscale scanning. We prove a uniform invariance principle for the resulting adjusted multiscale statistic (AMS), which is widely applicable and provides a computationally feasible way to simulate asymptotically correct critical values. We illustrate the implications of our theoretical results in a simulation study and in a real data example from super-resolution STED microscopy. This allows us to identify interesting regions inside a specimen in a pre-scan with controlled family-wise error rate.

stat.AP

Quick Adaptive Ternary Segmentation: An Efficient Decoding Procedure For Hidden Markov Models

Hidden Markov models (HMMs) are characterized by an unobservable Markov chain and an observable process -- a noisy version of the hidden chain. Decoding the original signal from the noisy observations is one of the main goals in nearly all HMM based data analyses. Existing decoding algorithms such as Viterbi and the pointwise maximum a posteriori (PMAP) algorithm have computational complexity at best linear in the length of the observed sequence, and sub-quadratic in the size of the state space of the hidden chain. We present Quick Adaptive Ternary Segmentation (QATS), a divide-and-conquer procedure with computational complexity polylogarithmic in the length of the sequence, and cubic in the size of the state space, hence particularly suited for large scale HMMs with relatively few states. It also suggests an effective way of data storage as specific cumulative sums. In essence, the estimated sequence of states sequentially maximizes local likelihood scores among all local paths with at most three segments, and is meanwhile admissible. The maximization is performed only approximately using an adaptive search procedure. Our simulations demonstrate the speedups offered by QATS in comparison to Viterbi and PMAP, along with a precision analysis. An implementation of QATS is in the R-package QATS on GitHub.

stat.ME

Empirical Optimal Transport under Estimated Costs: Distributional Limits and Statistical Applications

Optimal transport (OT) based data analysis is often faced with the issue that the underlying cost function is (partially) unknown. This paper is concerned with the derivation of distributional limits for the empirical OT value when the cost function and the measures are estimated from data. For statistical inference purposes, but also from the viewpoint of a stability analysis, understanding the fluctuation of such quantities is paramount. Our results find direct application in the problem of goodness-of-fit testing for group families, in machine learning applications where invariant transport costs arise, in the problem of estimating the distance between mixtures of distributions, and for the analysis of empirical sliced OT quantities. The established distributional limits assume either weak convergence of the cost process in uniform norm or that the cost is determined by an optimization problem of the OT value over a fixed parameter space. For the first setting we rely on careful lower and upper bounds for the OT value in terms of the measures and the cost in conjunction with a Skorokhod representation. The second setting is based on a functional delta method for the OT value process over the parameter space. The proof techniques might be of independent interest.

math.ST

Unbalanced Kantorovich-Rubinstein distance, plan, and barycenter on finite spaces: A statistical perspective

We analyze statistical properties of plug-in estimators for unbalanced optimal transport quantities between finitely supported measures in different prototypical sampling models. Specifically, our main results provide non-asymptotic bounds on the expected error of empirical Kantorovich-Rubinstein (KR) distance, plans, and barycenters for mass penalty parameter $C>0$. The impact of the mass penalty parameter $C$ is studied in detail. Based on this analysis, we mathematically justify randomized computational schemes for KR quantities which can be used for fast approximate computations in combination with any exact solver. Using synthetic and real datasets, we empirically analyze the behavior of the expected errors in simulation studies and illustrate the validity of our theoretical bounds.

stat.ME