Search arXiv⌕ Search

arXiv · 1009.0169

Quantitative test of the barrier nucleosome model for statistical positioning of nucleosomes up- and downstream of transcription start sites

Abstract

The positions of nucleosomes in eukaryotic genomes determine which parts of the DNA sequence are readily accessible for regulatory proteins and which are not. Genome-wide maps of nucleosome positions have revealed a salient pattern around transcription start sites, involving a nucleosome-free region (NFR) flanked by a pronounced periodic pattern in the average nucleosome density. While the periodic pattern clearly reflects well-positioned nucleosomes, the positioning mechanism is less clear. A recent experimental study by Mavrich et al. argued that the pattern observed in S. cerevisiae is qualitatively consistent with a `barrier nucleosome model', in which the oscillatory pattern is created by the statistical positioning mechanism of Kornberg and Stryer. On the other hand, there is clear evidence for intrinsic sequence preferences of nucleosomes, and it is unclear to what extent these sequence preferences affect the observed pattern. To test the barrier nucleosome model, we quantitatively analyze yeast nucleosome positioning data both up- and downstream from NFRs. Our analysis is based on the Tonks model of statistical physics which quantifies the interplay between the excluded-volume interaction of nucleosomes and their positional entropy. We find that although the typical patterns on the two sides of the NFR are different, they are both quantitatively described by the same physical model, with the same parameters, but different boundary conditions. The inferred boundary conditions suggest that the first nucleosome downstream from the NFR (the +1 nucleosome) is typically directly positioned while the first nucleosome upstream is statistically positioned via a nucleosome-repelling DNA region. These boundary conditions, which can be locally encoded into the genome sequence, significantly shape the statistical distribution of nucleosomes over a range of up to ~1000 bp to each side.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wolfram Mobius, Ulrich Gerland. 2010-09-01. Quantitative test of the barrier nucleosome model for statistical positioning of nucleosomes up- and downstream of transcription start sites. https://doi.org/10.1371/journal.pcbi.1000891

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

BreCol: Benchmarking Classical and Deep-Learning Methods for Microbiome-Based Cancer Detection

DNA sequencing of the gut microbial community shows promise for cancer detection, but questions remain about the generalizability of results across studies. We propose BreCol, a benchmark of 2,040 16S rRNA gene sequencing runs across 26 studies spanning breast cancer, colorectal cancer, and healthy cohorts. Train-test splits are made within pre-2023 studies, while holdout evaluation uses studies from 2023 onward, reflecting temporal separation from training data. Classical models reach test/holdout AUCs of 0.77/0.60 for cancer diagnosis and 1.00/0.83 for cancer type prediction. We train the models on both cancer types simultaneously and find that colorectal cancer is often easier to detect than breast cancer. We also evaluate two deep learning models: HyenaDNA, a long-range sequence model that pools hidden states for classification, and SetBERT, a transformer that produces contextualized embeddings over sets of reads. Both deep learning models underperform the best classical methods on holdout data, though tuning training set size and the classification head yields modest gains. Our classical pipeline uses unsupervised clustering to derive features from tetramer frequencies, preserving within-run compositional signal and achieving near state-of-the-art performance without relying on taxonomic assignments. BreCol data and associated code are publicly available.

q-bio.QM↗

Topological Inference for Organoids

The reproducibility of organ morphology and the extent to which computational models can predict morphogenesis remain difficult to quantify, particularly for organs with complex networks of fluid-filled lumina. Here, we combine Topological Data Analysis (TDA), biophysical simulation, and Bayesian inference to study lumen morphogenesis in pancreatic organoids. Lumen formation is governed by physical processes that are challenging to measure directly, including cell proliferation and luminal osmotic pressure. We simulate organoid development using a phase-field model and address the inverse problem of inferring these parameters from either time-lapse images or single morphological snapshots. Since lumen architectures vary substantially in size, structure, and connectivity, conventional geometric descriptors provide only a partial representation of their morphology. We therefore represent each organoid using SampEuler, a topological descriptor derived from the Euler Characteristic Transform (ECT). We first show that SampEuler captures morphological information encoded by established morphometrics. We then perform parameter inference using an approximate Bayesian computation (ABC) rejection framework with the SampEuler Wasserstein distance. Using synthetic organoids with known ground-truth parameters, our approach accurately recovers the osmotic pressure and the proliferation rate while revealing a compensatory trade-off between the two processes. Applied to experimental data from ten pancreatic organoids, the inferred posterior distributions are consistent with biological expectations. Together, these results establish a non-destructive, image-based pipeline for estimating otherwise inaccessible physical parameters governing lumen formation and highlight the potential of topological representations for linking complex biological morphology to mechanistic models.

q-bio.QM↗

Circadian Derived Features for Early Discrimination Across Insomnia Severity Levels: At Least 8 Weeks of Monitoring Are Needed for Clinically Meaningful Assessment

Background: Wearable devices provide continuous, objective measures of daily activity and offer promise for assessing sleep disorders. However, the minimum monitoring duration needed to differentiate insomnia severity remains unclear. We investigated when wearable-derived behavioral features become informative for distinguishing Insomnia Severity Index (ISI) categories and examined the contribution of Activity Count (AC) and circadian-derived features. Methods: We analyzed wearable data from 2,305 participants in the Advancing Understanding of Recovery after Trauma (AURORA) study. Separate binary classifiers were developed for four ISI categories across six follow-up periods using AC and circadian-derived feature sets. Performance was evaluated using accuracy, F1-score, precision, recall, and AUROC. Results: Classification improved with longer monitoring. Across ISI categories, approximately eight weeks was the earliest time point at which wearable-derived features consistently achieved informative discrimination, with modest improvements thereafter. Participants without clinically significant insomnia were easiest to identify, reaching an AUROC of 0.693. Circadian-derived features performed comparably to, and in several cases better than, AC features, suggesting that the temporal organization of daily activity provides information beyond overall activity volume. Conclusions: Approximately eight weeks of longitudinal wearable monitoring may represent a practical minimum for differentiating ISI-defined insomnia categories. Longer monitoring provided only incremental improvements. Circadian behavioral features show promise as digital biomarkers for objective insomnia assessment.

q-bio.QM↗