Search arXivSearch

arXiv · 1408.4224

CAPRI: Efficient Inference of Cancer Progression Models from Cross-sectional Data

Abstract

We devise a novel inference algorithm to effectively solve the cancer progression model reconstruction problem. Our empirical analysis of the accuracy and convergence rate of our algorithm, CAncer PRogression Inference (CAPRI), shows that it outperforms the state-of-the-art algorithms addressing similar problems. Motivation: Several cancer-related genomic data have become available (e.g., The Cancer Genome Atlas, TCGA) typically involving hundreds of patients. At present, most of these data are aggregated in a cross-sectional fashion providing all measurements at the time of diagnosis.Our goal is to infer cancer progression models from such data. These models are represented as directed acyclic graphs (DAGs) of collections of selectivity relations, where a mutation in a gene A selects for a later mutation in a gene B. Gaining insight into the structure of such progressions has the potential to improve both the stratification of patients and personalized therapy choices. Results: The CAPRI algorithm relies on a scoring method based on a probabilistic theory developed by Suppes, coupled with bootstrap and maximum likelihood inference. The resulting algorithm is efficient, achieves high accuracy, and has good complexity, also, in terms of convergence properties. CAPRI performs especially well in the presence of noise in the data, and with limited sample sizes. Moreover CAPRI, in contrast to other approaches, robustly reconstructs different types of confluent trajectories despite irregularities in the data.We also report on an ongoing investigation using CAPRI to study atypical Chronic Myeloid Leukemia, in which we uncovered non trivial selectivity relations and exclusivity patterns among key genomic events.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Daniele Ramazzotti, Giulio Caravagna, Loes Olde Loohuis, Alex Graudenzi, Ilya Korsunsky, Giancarlo Mauri, Marco Antoniotti, Bud Mishra. 2015-05-07. CAPRI: Efficient Inference of Cancer Progression Models from Cross-sectional Data. https://doi.org/10.1093/bioinformatics%2Fbtv296

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Functional independent component analysis by choice of norm: a framework for near-perfect classification

We develop a theory for functional independent component analysis in an infinite-dimensional framework using Sobolev spaces that accommodate smoother functions. The notion of penalized kurtosis is introduced motivated by Silverman's method for smoothing principal components. This approach allows for a classical definition of independent components obtained via projection onto the eigenfunctions of a smoothed kurtosis operator mapping a whitened functional random variable. We discuss the theoretical properties of this operator in relation to a generalized Fisher discriminant function and the relationship it entails with the Feldman-Hájek dichotomy for Gaussian measures, both of which are critical to the principles of functional classification. The proposed estimators are a particularly competitive alternative in binary classification of functional data and can eventually achieve the so-called near-perfect classification, which is a genuine phenomenon of high-dimensional data. Our methods are illustrated through simulations, various real datasets, and used to model electroencephalographic biomarkers for the diagnosis of depressive disorder.

math.ST

Trace-Class Results for MCMC Algorithms for Student-$t$ Regression Models

In this paper, we consider MCMC algorithms for Student-$t$ regression models. In three cases, we investigate the efficiency of Markov chains based on the algorithms in terms of whether trace-class results hold or not. First, we consider the case where the parameters follow a matrix-normal-inverse-Wishart distribution and show that the Markov operator associated with a standard data augmentation algorithm is trace-class. Second, we consider the case of an improper prior and univariate outcomes. In this case, the standard Markov operator is not trace-class but the Markov operator associated with a collapsed Gibbs algorithm is trace-class. Third, we consider the case of an improper prior and multivariate outcomes. We obtain a trace-class result for a parameter expanded data augmentation algorithm which is based on a univariate working parameter. Finally, we consider the problem of numerially estimating a convergence rate of the trace-class Markov operator in the second case.

math.ST

The Manifold Hypothesis under Unknown Gaussian Noise:Conditional Certificates and Consistent Dimension Estimation

We study what noisy data can establish about the Manifold Hypothesis under explicit identification and regularity conditions. A population residual certificate combines independent-view localization, Gaussian concentration, membership uncertainty, and population transfer. Existing rectifiability criteria then yield a covered-scale consequence. For a local smooth manifold with positive Hölder density, the actual-ball covariance limit identifies the spectral crossing with geometric dimension. We prove almost-sure eventual recovery under repeated observations. Reusing accurate localization averages improves the sufficient point-sample condition from $Nr^{d+4}\gg\log N$ to $Nr^d\gg\log N$, with replication $kr^2\gg\log N$. A two-mass certificate controls incorrect geometric-dimension emissions under declared class bounds. For single observations with unknown Gaussian noise, affine-support or known coordinate-bound restrictions provide noise intervals and consistent Gaussian correlation-dimension estimators. Ahlfors regularity identifies this exponent with Hausdorff dimension and with the geometric dimension of a homogeneous smooth class. Exact Cantor calculations delineate the limits of integer spectral counts and adjacent-radius slopes. We credit established local PCA, rectifiability, concentration, binomial inference, and deconvolution results before specifying our constructions. Reproducible experiments distinguish point estimation, finite-scale coverage, and certificate emission.

math.ST