Search arXivSearch

arXiv · 1803.09021

On Large-Scale Graph Generation with Validation of Diverse Triangle Statistics at Edges and Vertices

Abstract

Researchers developing implementations of distributed graph analytic algorithms require graph generators that yield graphs sharing the challenging characteristics of real-world graphs (small-world, scale-free, heavy-tailed degree distribution) with efficiently calculable ground-truth solutions to the desired output. Reproducibility for current generators used in benchmarking are somewhat lacking in this respect due to their randomness: the output of a desired graph analytic can only be compared to expected values and not exact ground truth. Nonstochastic Kronecker product graphs meet these design criteria for several graph analytics. Here we show that many flavors of triangle participation can be cheaply calculated while generating a Kronecker product graph. Given two medium-sized scale-free graphs with adjacency matrices $A$ and $B$, their Kronecker product graph has adjacency matrix $C = A \otimes B$. Such graphs are highly compressible: $|{\cal E}|$ edges are represented in ${\cal O}(|{\cal E}|^{1/2})$ memory and can be built in a distributed setting from small data structures, making them easy to share in compressed form. Many interesting graph calculations have worst-case complexity bounds ${\cal O}(|{\cal E}|^p)$ and often these are reduced to ${\cal O}(|{\cal E}|^{p/2})$ for Kronecker product graphs, when a Kronecker formula can be derived yielding the sought calculation on $C$ in terms of related calculations on $A$ and $B$. We focus on deriving formulas for triangle participation at vertices, ${\bf t}_C$, a vector storing the number of triangles that every vertex is involved in, and triangle participation at edges, $Δ_C$, a sparse matrix storing the number of triangles at every edge.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Geoffrey Sanders, Roger Pearce, Timothy La Fond, Jeremy Kepner. 2018-03-24. On Large-Scale Graph Generation with Validation of Diverse Triangle Statistics at Edges and Vertices. https://doi.org/10.1109/ipdpsw.2018.00056

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Factorisability of Low Dimensional Non-Negative Integer Matrices

We consider the problem of determining if a given two-dimensional nonnegative integer matrix $M$ is the product of two such matrices, excluding trivial units. A matrix $M$ with no such factorisation is called prime and therefore belongs to the minimal (infinite rank) generator of $2 \times 2$ matrices over the natural numbers, otherwise it is called composite. We also consider the problem of finding a (non-unique) factorisation of a composite matrix. Our results have applications in computational group theory and the theory of codes, where such matrices are called incidence matrices. We analyse the complexity of primality and finding a factorisation for a composite matrix, providing a first efficient algorithm.

cs.DM

Three Hardness Results for Graph Similarity Problems

Notions of graph similarity provide alternative perspective on the graph isomorphism problem and vice-versa. In this paper, we consider measures of similarity arising from mismatch norms as studied in Gervens and Grohe: the edit distance $δ_{\mathcal{E}}$, and the metrics arising from $\ell_p$-operator norms, which we denote by $δ_p$ and $δ_{|p|}$. We address the following question: can these measures of similarity be used to design polynomial-time approximation algorithms for graph isomorphism? We show that computing an optimal value of $δ_{\mathcal{E}}$ is \NP-hard on pairs of graphs with the same number of edges. In addition, we show that computing optimal values of $δ_p$ and $δ_{|p|}$ is \NP-hard even on pairs of $1$-planar graphs with the same degree sequence and bounded degree. These two results improve on previous known ones, which did not examine the restricted case where the pairs of graphs are required to have the same number of edges. Finally, we study similarity problems on strongly regular graphs and prove some near optimal inequalities with interesting consequences on the computational complexity of graph and group isomorphism.

cs.DM