Search arXivSearch

arXiv subjects

David Uminsky

Publications and source records attributed to David Uminsky.

13 recordsLinked to original sources

LUM\'AWIG: An Efficient Algorithm for Dimension Zero Bottleneck Distance Computation in Topological Data Analysis

Stability of persistence diagrams under slight perturbations is a key characteristic behind the validity and growing popularity of topological data analysis in exploring real-world data. Central to this stability is the use of Bottleneck distance which entails matching points between diagrams. Use of this metric in practical studies has, however, been few and sparingly because of the computational obstruction, especially in dimension zero where the computational cost explodes with the growth of data size. We present LUM\'AWIG, a novel efficient algorithm to compute dimension zero bottleneck distance between two persistent diagrams which runs significantly faster and provides significantly sharper approximates with respect to the output of the original algorithm than any other available algorithm. We bypass the overwhelming matching problem in previous implementations of the bottleneck distance, and prove that the zero dimensional bottleneck distance can be recovered from a very small number of matching cases. We show that LUM\'AWIG generally enjoys linear complexity as shown by empirical tests. We also present an application that leverages dimension zero persistence diagrams and the bottleneck distance to produce features for classification tasks.

cs.CG

Identifying group contributions in NBA lineups with spectral analysis

We address the question of how to quantify the contributions of groups of players to team success. Our approach is based on spectral analysis, a technique from algebraic signal processing, which has several appealing features. First, our analysis decomposes the team success signal into components that are naturally understood as the contributions of player groups of a given size: individuals, pairs, triples, fours, and full five-player lineups. Secondly, the decomposition is orthogonal so that contributions of a player group can be thought of as pure: Contributions attributed to a group of three, for example, have been separated from the lower-order contributions of constituent pairs and individuals. We present detailed a spectral analysis using NBA play-by-play data and show how this can be a practical tool in understanding lineup composition and utilization.

stat.AP

The Problem with Metrics is a Fundamental Problem for AI

Optimizing a given metric is a central aspect of most current AI approaches, yet overemphasizing metrics leads to manipulation, gaming, a myopic focus on short-term goals, and other unexpected negative consequences. This poses a fundamental contradiction for AI development. Through a series of real-world case studies, we look at various aspects of where metrics go wrong in practice and aspects of how our online environment and current business practices are exacerbating these failures. Finally, we propose a framework towards mitigating the harms caused by overemphasis of metrics within AI by: (1) using a slate of metrics to get a fuller and more nuanced picture, (2) combining metrics with qualitative accounts, and (3) involving a range of stakeholders, including those who will be most impacted.

cs.CY

Classification of Single-lead Electrocardiograms: TDA Informed Machine Learning

Atrial Fibrillation is a heart condition characterized by erratic heart rhythms caused by chaotic propagation of electrical impulses in the atria, leading to numerous health complications. State-of-the-art models employ complex algorithms that extract expert-informed features to improve diagnosis. In this note, we demonstrate how topological features can be used to help accurately classify single lead electrocardiograms. Via delay embeddings, we map electrocardiograms onto high-dimensional point-clouds that convert periodic signals to algebraically computable topological signatures. We derive features from persistent signatures, input them to a simple machine learning algorithm, and benchmark its performance against winning entries in the 2017 Physionet Computing in Cardiology Challenge.

q-bio.QM

A Non-iterative Parallelizable Eigenbasis Algorithm for Johnson Graphs

We present a new $O(k^2 \binom{n}{k}^2)$ method for generating an orthogonal basis of eigenvectors for the Johnson graph $J(n,k)$. Unlike standard methods for computing a full eigenbasis of sparse symmetric matrices, the algorithm presented here is non-iterative, and produces exact results under an infinite-precision computation model. In addition, our method is highly parallelizable; given access to unlimited parallel processors, the eigenbasis can be constructed in only $O(n)$ time given n and k. We also present an algorithm for computing projections onto the eigenspaces of $J(n,k)$ in parallel time $O(n)$.

cs.DS

Multiclass Total Variation Clustering

Ideas from the image processing literature have recently motivated a new set of clustering algorithms that rely on the concept of total variation. While these algorithms perform well for bi-partitioning tasks, their recursive extensions yield unimpressive results for multiclass clustering tasks. This paper presents a general framework for multiclass total variation clustering that does not rely on recursion. The results greatly outperform previous total variation algorithms and compare well with state-of-the-art NMF approaches.

stat.ML

An Adaptive Total Variation Algorithm for Computing the Balanced Cut of a Graph

We propose an adaptive version of the total variation algorithm proposed in [3] for computing the balanced cut of a graph. The algorithm from [3] used a sequence of inner total variation minimizations to guarantee descent of the balanced cut energy as well as convergence of the algorithm. In practice the total variation minimization step is never solved exactly. Instead, an accuracy parameter is specified and the total variation minimization terminates once this level of accuracy is reached. The choice of this parameter can vastly impact both the computational time of the overall algorithm as well as the accuracy of the result. Moreover, since the total variation minimization step is not solved exactly, the algorithm is not guarantied to be monotonic. In the present work we introduce a new adaptive stopping condition for the total variation minimization that guarantees monotonicity. This results in an algorithm that is actually monotonic in practice and is also significantly faster than previous, non-adaptive algorithms.

math.OC

Convergence of a Steepest Descent Algorithm for Ratio Cut Clustering

Unsupervised clustering of scattered, noisy and high-dimensional data points is an important and difficult problem. Tight continuous relaxations of balanced cut problems have recently been shown to provide excellent clustering results. In this paper, we present an explicit-implicit gradient flow scheme for the relaxed ratio cut problem, and prove that the algorithm converges to a critical point of the energy. We also show the efficiency of the proposed algorithm on the two moons dataset.

math.OC

A multi-moment vortex method for 2D viscous fluids

In this paper we introduce simplified, combinatorially exact formulas that arise in the vortex interaction model found in (Nagem, et al., SIAM J. Appl. Dyn. Syst. 2009). These combinatorial formulas allow for the efficient implementation and development of a new multi-moment vortex method (MMVM) using a Hermite expansion to simulate 2D vorticity. The method naturally allows the particles to deform and become highly anisotropic as they evolve without the added cost of computing the non-local Biot-Savart integral. We present three examples using MMVM. We first focus our attention on the implementation of a single particle, large number of Hermite moments case, in the context of quadrupole perturbations of the Lamb-Oseen vortex. At smaller perturbation values, we show the method captures the shear diffusion mechanism and the rapid relaxation (on $Re^{1/3}$ time scale) to an axisymmetric state. We then present two more examples of the full multi-moment vortex method and discuss the results in the context of classic vortex methods. We perform spatial convergence studies of the single-particle method and show that the method exhibits exponential convergence. Lastly, we numerically investigate the spatial accuracy improvement from the inclusion of higher Hermite moments in the full MMVM.

physics.flu-dyn

The QCD beta-function from global solutions to Dyson-Schwinger equations

We study quantum chromodynamics from the viewpoint of untruncated Dyson-Schwinger equations turned to an ordinary differential equation for the gluon anomalous dimension. This nonlinear equation is parameterized by a function P(x) which is unknown beyond perturbation theory. Still, very mild assumptions on P(x) lead to stringent restrictions for possible solutions to Dyson-Schwinger equations. We establish that the theory must have asymptotic freedom beyond perturbation theory and also investigate the low energy regime and the possibility for a mass gap in the asymptotically free theory.

hep-th

The QED beta-function from global solutions to Dyson-Schwinger equations

We discuss the structure of beta functions as determined by the recursive nature of Dyson--Schwinger equations turned into an analysis of ordinary differential equations, with particular emphasis given to quantum electrodynamics. In particular we determine when a separatrix for solutions to such ODEs exists and clarify the existence of Landau poles beyond perturbation theory. Both are determined in terms of explicit conditions on the asymptotics for the growth of skeleton graphs.

hep-th

Generalized Helmholtz-Kirchhoff model for two dimensional distributed vortex motion

The two-dimensional Navier-Stokes equations are rewritten as a system of coupled nonlinear ordinary differential equations. These equations describe the evolution of the moments of an expansion of the vorticity with respect to Hermite functions and of the centers of vorticity concentrations. We prove the convergence of this expansion and show that in the zero viscosity and zero core size limit we formally recover the Helmholtz-Kirchhoff model for the evolution of point-vortices. The present expansion systematically incorporates the effects of both viscosity and finite vortex core size. We also show that a low-order truncation of our expansion leads to the representation of the flow as a system of interacting Gaussian (i.e. Oseen) vortices which previous experimental work has shown to be an accurate approximation to many important physical flows [9].

math.DS

Unbounded regions of Infinitely Logconcave Sequences

We study the properties of a logconcavity operator on a symmetric, unimodal subset of finite sequences. In doing so we are able to prove that there is a large unbounded region in this subset that is $\infty$-logconcave. This problem was motivated by the conjecture of Moll and Boros in that the binomial coefficients are $\infty$-logconcave.

math.CO