Search arXiv⌕ Search

arXiv · q-bio/0702029

Graph animals, subgraph sampling and motif search in large networks

Abstract

We generalize a sampling algorithm for lattice animals (connected clusters on a regular lattice) to a Monte Carlo algorithm for `graph animals', i.e. connected subgraphs in arbitrary networks. As with the algorithm in [N. Kashtan et al., Bioinformatics 20, 1746 (2004)], it provides a weighted sample, but the computation of the weights is much faster (linear in the size of subgraphs, instead of super-exponential). This allows subgraphs with up to ten or more nodes to be sampled with very high statistics, from arbitrarily large networks. Using this together with a heuristic algorithm for rapidly classifying isomorphic graphs, we present results for two protein interaction networks obtained using the TAP high throughput method: one of Escherichia coli with 230 nodes and 695 links, and one for yeast (Saccharomyces cerevisiae) with roughly ten times more nodes and links. We find in both cases that most connected subgraphs are strong motifs (Z-scores >10) or anti-motifs (Z-scores <-10) when the null model is the ensemble of networks with fixed degree sequence. Strong differences appear between the two networks, with dominant motifs in E. coli being (nearly) bipartite graphs and having many pairs of nodes which connect to the same neighbors, while dominant motifs in yeast tend towards completeness or contain large cliques. We also explore a number of methods that do not rely on measurements of Z-scores or comparisons with null models. For instance, we discuss the influence of specific complexes like the 26S proteasome in yeast, where a small number of complexes dominate the $k$-cores with large k and have a decisive effect on the strongest motifs with 6 to 8 nodes. We also present Zipf plots of counts versus rank. They show broad distributions that are not power laws, in contrast to the case when disconnected subgraphs are included.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kim Baskerville, Peter Grassberger, Maya Paczuski. 2007-06-22. Graph animals, subgraph sampling and motif search in large networks. https://doi.org/10.1103/physreve.76.036107

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Minimality in Reflexive and Stoichiometric Autocatalysis

Autocatalysis, the ability of a chemical subsystem to sustain its own constituents when supplied with sufficient food molecules, has been closely related to the origin of life on Earth. Emerging from Wilhelm Ostwald's considerations about an explicit autocatalytic reaction, different notions of autocatalysis have been developed over the years. The two most prominent are reflexively autocatalytic F-generated sets (RAFs) and stoichiometric autocatalysis. After having shown that each RAFs is, under reasonable conditions, in general stoichiometrically autocatalytic, we examine here the relationship between the two notions of minimality: irreducible RAFs and autocatalytic cores. To this end, we overcome the obstacle that RAFs and stoichiometric autocatalysis have been formalized in distinct systems of chemical reactions, i.e., catalytic reaction systems (CRS) and chemical reaction networks (CRNs), respectively. We show that reactions in a CRS constitute equivalence classes of reactions of the corresponding CRNs w.r.t. their specific catalyzations. Using the fact that CRN and CRS can be canonically identified whenever each CRS reaction is associated with a single catalyzation, we demonstrate that the Kőnig graph of a monocatalyzed, irreducible RAF is composed of strong blocks devoid of food and waste species that are separated by reaction vertices, each of which contains an autocatalytic core. In fact, a single irreducible RAF can, in general, contain exponentially many autocatalytic cores.

q-bio.MN↗

Implementation of Linear Regression and Linear Interpolation using Reaction Networks

Statistical inference is a fundamental component of data science. In this work, we focus on two classical inference techniques: regression and interpolation. We propose a reaction-network-based framework for implementing linear regression, including both univariate and multivariate settings, as well as linear interpolation. Our approach encodes the outputs of these inference techniques in the steady-state concentrations of species within the reaction network. A key ingredient of the construction is a novel generalized division module capable of handling division involving negative numbers. We validate the proposed framework through in silico implementations on standard synthetic datasets and obtain the expected regression and interpolation outputs.

q-bio.MN↗

Mapping disease-regulatory flux through eQTL-based causal gene networks: a complex-trait framework applied to coronary artery disease

A substantial proportion of inherited susceptibility to common diseases is mediated through tissue-specific regulatory variation. The omnigenic model hypothesizes that much of this risk arises from distant (trans) regulatory effects that are propagated via gene-regulatory networks (GRNs) from numerous regulatory genes onto a relatively small set of core genes. Two major challenges have hindered an empirical assessment of this model. First, transcriptome-wide association studies primarily assess gene expression that is regulated locally (in cis), and thus miss trans-acting effects. Second, trans regulation has so far been examined only in aggregate, without identifying which regulatory genes transmit disease-associated signals to which target genes. Here, we present a framework that decomposes each gene's disease association into a local cis component and a set of trans components attributable to its upstream regulators, propagated along a directed causal GRN. For each gene, we estimate sparse Bayesian expression models and we assess model performance using out-of-sample prediction. Propagating trans signals through the inferred network yields a disease-regulatory flux map: a directed, signed representation that quantifies the contribution of each regulator to the disease association of each target gene. Applying this framework to seven tissues relevant to coronary artery disease, we demonstrate how genetic variation flows through the GRN to affect disease risk. Outgoing regulatory influences from individual genes are directionally heterogeneous, whereas disease-associated genes tend to integrate convergent input from multiple regulators; these convergent targets are enriched for cardiovascular-related biological processes. Consequently, each gene's disease association can be reinterpreted as a detailed allocation of the disease signal among its contributing regulators.

q-bio.MN↗