Search arXivSearch

arXiv subjects

Leonardo Varuzza

Publications and source records attributed to Leonardo Varuzza.

2 recordsLinked to original sources

Significance tests for comparing digital gene expression profiles

Most of the statistical tests currently used to detect differentially expressed genes are based on asymptotic results, and perform poorly for low expression tags. Another problem is the common use of a single canonical cutoff for the significance level (p-value) of all the tags, without taking into consideration the type II error and the highly variable character of the sample size of the tags. This work reports the development of two significance tests for the comparison of digital expression profiles, based on frequentist and Bayesian points of view, respectively. Both tests are exact, and do not use any asymptotic considerations, thus producing more correct results for low frequency tags than the chi-square test. The frequentist test uses a tag-customized critical level which minimizes a linear combination of type I and type II errors. A comparison of the Bayesian and the frequentist tests revealed that they are linked by a Beta distribution function. These tests can be used alone or in conjunction, and represent an improvement over the currently available methods for comparing digital profiles.

q-bio.GN

Simcluster: clustering enumeration gene expression data on the simplex space

Transcript enumeration methods such as SAGE, MPSS, and sequencing-by-synthesis EST ``digital northern'', are important high-throughput techniques for digital gene expression measurement. As other counting or voting processes, these measurements constitute compositional data exhibiting properties particular to the simplex space where the summation of the components is constrained. These properties are not present on regular Euclidean spaces, on which hybridization-based microarray data is often modeled. Therefore, pattern recognition methods commonly used for microarray data analysis may be non-informative for the data generated by transcript enumeration techniques since they ignore certain fundamental properties of this space. Here we present a software tool, Simcluster, designed to perform clustering analysis for data on the simplex space. We present Simcluster as a stand-alone command-line C package and as a user-friendly on-line tool. Both versions are available at: http://xerad.systemsbiology.net/simcluster/ . Simcluster is designed in accordance with a well-established mathematical framework for compositional data analysis, which provides principled procedures for dealing with the simplex space, and is thus applicable in a number of contexts, including enumeration-based gene expression data.

q-bio.QM