Search arXivSearch

arXiv · 1708.04024

Create, run, share, publish, and reference your LC-MS, FIA-MS, GC-MS, and NMR data analysis workflows with the Workflow4Metabolomics 3.0 Galaxy online infrastructure for metabolomics

Abstract

Metabolomics is a key approach in modern functional genomics and systems biology. Due to the complexity of metabolomics data, the variety of experimental designs, and the variety of existing bioinformatics tools, providing experimenters with a simple and efficient resource to conduct comprehensive and rigorous analysis of their data is of utmost importance. In 2014, we launched the Workflow4Metabolomics (W4M, http://workflow4metabolomics.org) online infrastructure for metabolomics built on the Galaxy environment, which offers user-friendly features to build and run data analysis workflows including preprocessing, statistical analysis, and annotation steps. Here we present the new W4M 3.0 release, which contains twice as many tools as the first version, and provides two features which are, to our knowledge, unique among online resources. First, data from the four major metabolomics technologies (i.e., LC-MS, FIA-MS, GC-MS, and NMR) can be analyzed on a single platform. By using three studies in human physiology, alga evolution, and animal toxicology, we demonstrate how the 40 available tools can be easily combined to address biological issues. Second, the full analysis (including the workflow, the parameter values, the input data and output results) can be referenced with a permanent digital object identifier (DOI). Publication of data analyses is of major importance for robust and reproducible science. Furthermore, the publicly shared workflows are of high-value for e-learning and training. The Workflow4Metabolomics 3.0 e-infrastructure thus not only offers a unique online environment for analysis of data from the main metabolomics technologies, but it is also the first reference repository for metabolomics workflows.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yann Guitton, Marie Tremblay-Franco, Gildas Le Corguillé, Jean-François Martin, Mélanie Pétéra, Pierrick Roger-Mele, Alexis Delabrière, Sophie Goulitquer, Misharl Monsoor, Christophe Duperier, Cécile Canlet, Remi Servien, Patrick Tardivel, Christophe Caron, Franck Giacomoni, Etienne Thévenot. 2017-08-14. Create, run, share, publish, and reference your LC-MS, FIA-MS, GC-MS, and NMR data analysis workflows with the Workflow4Metabolomics 3.0 Galaxy online infrastructure for metabolomics. https://doi.org/10.1016/j.biocel.2017.07.002

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Triplication: an important component of the modern scientific method

A scientific-study protocol (defined) is designed to deliver results from which inductive inference is allowed. In the nineteenth century, triplication was introduced into the plant sciences and Fisher's p<0.05 rule (1925) incorporated into triple-result protocols designed to counter random/systematic errors which contribute to real-world variability. The aims of the present study were to: (1) classify replication protocols; (2) assess their prevalence in plant-science studies (published during one twenty-first-century year; for defined variable construct); (3) explore triplication rationale. Methods: a plant-sciences protocol-prevalence report was produced; experimental/associational-study proportions analyzed; and real-world-data proxies used to show confidence-interval-width patterns with increasing replicate number. Results: 25% plant-science studies analyzed showed triplication, including 11% triple-result protocols (including greater replicate numbers: 48%;17%, respectively). Theoretical considerations indicated that even if systematic errors predominate, (previously-known) square-root rules sometimes apply, contributing to triplication importance (exemplified by real-world-data proxies). Conclusions: The defined protocols, with minor modifications, should provide the means for assessment of most sciences. Triplication was extensively applied in studies analysed and there are strong methodological reasons why triplication, rather than duplication/quadruplication, is the appropriate standard: triple-result protocols: (a) effectively reduce false positives to acceptable levels; (b) give qualitatively-different information (shape) from duplication; (c) have a large efficiency advantage (concerning confidence-interval widths) over quadruplication. The application of batch replication is not, primarily, a statistical problem and cannot effectively be replaced by simulation.

q-bio.QM

Retracing the Process of Translation: Proteome-wide mapping of stable transcriptomic predictors of protein abundance in cancer cell lines

Understanding the relationship between gene expression and protein abundance is central to molecular and systems biology. While gene expression reflects transcriptional activity, proteins are the functional molecules that determine cellular phenotypes. However, numerous post-transcriptional and translational regulatory layers complicate this relationship, and prior studies have reported only weak to moderate correlations between RNA and protein levels. Predicting protein abundance from transcriptomic data remains challenging, but it is a valuable goal for biological insight, especially when proteomic data is limited or unavailable. In this study, we applied a large-scale, Ridge regression-based feature selection strategy to identify predictive gene expression features for each of 8,423 proteins across 940 cancer cell lines. To our knowledge, this is the first work to perform such comprehensive protein-wise feature selection at this scale. Our analysis revealed both globally predictive and context-specific gene features. These included biologically meaningful modules such as immune-related genes, HOX transcription factor targets, and cytoskeletal components. The models identified stable candidate gene-protein associations that remained interpretable at the level of individual proteins and recurrent transcriptomic predictor patterns. Our approach enables interpretable modeling of protein expression from transcriptomic data and provides insight into transcriptomic features associated with protein abundance. This framework may support hypothesis generation, protein imputation in incomplete datasets, and deeper understanding of post-transcriptional regulation in cancer biology.

q-bio.QM

Surf_2_Volume: a workflow for converting CIFTI parcellations to NIfTI volume space

Many neuroimaging programs require volumetric NIfTI files and cannot directly use parcellations stored in CIFTI format. We developed Surf_2_Volume, a workflow that uses Connectome Workbench, FreeSurfer, AFNI/SUMA, neuromaps, and Python to convert categorical CIFTI parcellations into NIfTI volumes. Cortical labels are transferred through fsaverage to a surface representation of the MNI152 template and assigned to voxels within a smoothed cortical ribbon mask. A threshold controls how much of this mask is included. We tested twelve thresholds with the Schaefer2018 atlas with 100 parcels and 7 networks, which is available as both an fsLR representation and a published FSL MNI152 1 mm volume. We compared each converted volume with the published reference and with standard Workbench ribbon and nearest-vertex mappings. At threshold 0.05, Surf_2_Volume had higher scores on the two measures of parcel overlap, Dice and Jaccard, than the best Workbench setting (0.729 vs. 0.709 and 0.580 vs. 0.556, respectively). It also left fewer reference voxels unlabeled (11.88% vs. 20.57%) and assigned fewer labels outside the reference (9.43% vs. 11.22%). Among the tested Surf_2_Volume thresholds, 0.05 also produced the lowest average parcel volume distortion. Surf_2_Volume provides a tunable and reproducible workflow for generating volumetric atlases from CIFTI parcellations when volume-based analyses are required.

q-bio.QM