Search arXiv⌕ Search

arXiv · 0709.4206

In silico network topology-based prediction of gene essentiality

Abstract

The identification of genes essential for survival is important for the understanding of the minimal requirements for cellular life and for drug design. As experimental studies with the purpose of building a catalog of essential genes for a given organism are time-consuming and laborious, a computational approach which could predict gene essentiality with high accuracy would be of great value. We present here a novel computational approach, called NTPGE (Network Topology-based Prediction of Gene Essentiality), that relies on network topology features of a gene to estimate its essentiality. The first step of NTPGE is to construct the integrated molecular network for a given organism comprising protein physical, metabolic and transcriptional regulation interactions. The second step consists in training a decision tree-based machine learning algorithm on known essential and non-essential genes of the organism of interest, considering as learning attributes the network topology information for each of these genes. Finally, the decision tree classifier generated is applied to the set of genes of this organism to estimate essentiality for each gene. We applied the NTPGE approach for discovering essential genes in Escherichia coli and then assessed its performance.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Joao Paulo Muller da Silva, Marcio Luis Acencio, Jose Carlos Merino Mombach, Renata Vieira, Jose Guliherme Camargo da Silva, Ney Lemke, Marialva Sinigaglia. 2007-09-26. In silico network topology-based prediction of gene essentiality. https://arxiv.org/abs/0709.4206

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Unified Unsupervised Framework for Genome-Wide Association Studies in Heterogeneous Populations

Genome-wide association studies (GWAS) have greatly advanced the discovery of genetic variants underlying complex traits and diseases. Yet in heterogeneous populations, existing GWAS strategies typically either pool all individuals under an assumption of population homogeneity or perform meta-analysis across predefined subgroups, both of which are limited when latent genetic heterogeneity attenuates subgroup-specific effects and masks true associations or subgroup labels are imprecise. Here we present UCALM, a unified unsupervised framework that infers genetically homogeneous subgroups directly from the data and integrates subgroup-specific GWAS with a novel layered meta-analysis method to capture both shared and subgroup-specific association signals. Through extensive simulations and analyses of large-scale human and livestock cohorts, including the UK Biobank ($n \approx 487{,}000$) and a heterogeneous pig cohort ($n \approx 85{,}000$), we demonstrate that UCALM substantially alleviated the mean genomic inflation across 24 UK Biobank traits to 1.17 compared with 1.43 for GLM and 1.37 for LDAK-KVIK, and further revealed 74 loci in the pig cohort that were previously obscured by conventional approaches. Our results establish a robust and broadly applicable strategy for association mapping in structured populations and improve the resolution of genetic signals across diverse species.

q-bio.GN↗

Tracing model-generated DNA with position-independent watermarking

Genomic language models can write synthetic DNA that carries no record of its origin. A generation-time watermark could provide a provenance signal, if a verifier can detect it using only the DNA sequence, a key, and published detector settings, without knowing where the generated region starts, which strand it lies on, or how it was divided into six-base tokens. We embed the SynthID tournament watermark in two genomic language models, Carbon and GENERator-v2. The verifier searches both strands, every start position, and four window lengths, and corrects its decision for the whole search. In a detection cohort of 1,544 held-out prompts, it found all 3,088 marked sequences per model, unedited and after one substituted, inserted, or deleted base. For ordinary sequences, the one-sided 95% upper confidence bound on the false-positive rate was at most 0.850%; one bound for the wrong-key control reached 1.015% (Carbon, after a deletion). In a development cohort, neither model showed a detectable change in likelihood or in predefined sequence measures. Under random edits placed without reference to the detector, detection stayed complete or nearly complete up to a 2% per-base edit rate. These results show statistical detection, not biological function, secret-key security, or robustness to an editor who sees the detector.

q-bio.GN↗

Generating eukaryotic reference genome assemblies: Earth BioGenome Project quality standards and recommendations

Reference genome assemblies are foundational resources across the biological sciences as they begin to expose the fundamental building blocks of each species and the molecular toolkits vital for adaptation and survival. High-quality genomes enable insights into evolution, facilitate ecosystem monitoring and protection and conservation of species, access to new biomaterials and biomedicine, advances in agri- and aquaculture, and support planetary health. The amount of information gathered from a species' genome assembly is directly dependent on its quality. However, different sequencing technologies and software solutions generate a wide range of quality outcomes. Here, the Earth BioGenome Project's Sequencing and Assembly committee, together with a large community of researchers performing eukaryote sequencing and assembly, outlines quality standards for a "reference" assembly and formulates recommendations to ensure these standards are met.

q-bio.GN↗