Search arXivSearch

arXiv · 1206.0773

Changepoint Detection over Graphs with the Spectral Scan Statistic

Abstract

We consider the change-point detection problem of deciding, based on noisy measurements, whether an unknown signal over a given graph is constant or is instead piecewise constant over two connected induced subgraphs of relatively low cut size. We analyze the corresponding generalized likelihood ratio (GLR) statistics and relate it to the problem of finding a sparsest cut in a graph. We develop a tractable relaxation of the GLR statistic based on the combinatorial Laplacian of the graph, which we call the spectral scan statistic, and analyze its properties. We show how its performance as a testing procedure depends directly on the spectrum of the graph, and use this result to explicitly derive its asymptotic properties on few significant graph topologies. Finally, we demonstrate both theoretically and by simulations that the spectral scan statistic can outperform naive testing procedures based on edge thresholding and $χ^2$ testing.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

James Sharpnack, Alessandro Rinaldo, Aarti Singh. 2012-06-04. Changepoint Detection over Graphs with the Spectral Scan Statistic. https://arxiv.org/abs/1206.0773

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A global spectral gap for Metropolis-adjusted Langevin algorithm with a uniformly randomized step size

Let $π(\mathrm{d} x)\propto e^{-U(x)}\,\mathrm{d} x$ on $\mathbb{R}^d$, where $U$ is continuously differentiable and $m$-strongly convex with a globally $L$-Lipschitz gradient, $0<m\leq L<\infty$, and $κ=L/m$. It is known that, under warm-start assumptions, fixed-step Metropolis-adjusted Langevin algorithm (MALA) with properly tuned step size has mixing time of order $κ\sqrt{d}$ up to logarithmic factors. By contrast, when the condition number is bounded away from one, no single fixed step size yields a matching spectral-gap lower bound of order $(κ\sqrt{d})^{-1}$ uniformly over this target class. We establish a global spectral-gap lower bound for MALA with a uniformly randomized step size. At each iteration, the algorithm draws $h$ uniformly from $(0,H)$ and performs one ordinary MALA transition. Choosing $H$ of order $$\frac{1}{L\sqrt{d[1+\log(d+1)+\logκ]}}$$ yields a right spectral-gap lower bound of order $$\frac{1}{κ\sqrt{d[1+\log(d+1)+\logκ]}},$$ uniformly over the target class. A key ingredient in the proof is a Cheeger-type inequality for aggregating estimates of the one-step flow of MALA out of measurable sets at various step-size scales. It allows the scale used to control the flow to depend on the set and avoids the additional loss that would result from first estimating the conductance and then applying the standard Cheeger inequality. This work was developed with substantial assistance from ChatGPT, which suggested the uniformly randomized-step approach, developed the principal proof arguments, and generated the simulation and Lean 4 code. The human author checked and verified the mathematical content and takes full responsibility for the results.

math.ST

Evidence thresholds of collapsed structured block models

A stochastic block model explains a network by splitting its nodes into communities and a matrix of tie probabilities between them. Collapsing those probabilities out scores each split by one number, its marginal likelihood, and every decision the model takes, how many communities, whether a weak one survives, what generates the ties inside it, is a comparison of two such numbers. The comparisons fail in known ways: weak structure is dropped, small communities are merged, a group of hubs is returned in place of the real groups. Each failure has a threshold, and the thresholds have not been computed. We compute them. We prove that the best single partition keeps a planted structure only above the first-moment bound, which sits above the Kesten-Stigum threshold for up to ten communities, while the evidence summed over all partitions switches at Kesten-Stigum itself: between the two, the posterior favours structure that every partition a sampler can report rejects. We prove three further thresholds, one for merging small communities, one for the density at which the representation alone decides, one for the degree prior that cuts a community along its degrees. The empirical finding confirms the theory: on planted partitions and on thirteen labelled networks, from the karate club to citation and co-purchase networks of up to 19,717 nodes, each threshold falls where it is predicted, against likelihood, spectral, penalised and nonparametric alternatives.

math.ST

On the consistency of the posterior distribution for nonlinear PDE parameter identification

In this work, we investigate the estimation of a parameter in PDEs using Bayesian procedures, and focus on posterior distributions constructed using Gaussian process priors, and its variational approximation. We establish contraction rates for the posterior distribution and the variational approximation in the regime of low-regularity parameters. Specifically, the ground truth is only assumed to belong to the support space of the prior rather than to its associated reproducing kernel Hilbert space. Also we derive contraction rates for the posterior distribution constructed using the randomly truncated Gaussian process prior. The analysis relies on a delicate approximation argument that approximates low-regularity ground truths by suitable elements in the reproducing kernel Hilbert space and balances various error sources. We illustrate the general theory on three nonlinear inverse problems for PDEs.

math.ST