Search arXiv⌕ Search

arXiv · 0706.3137

Asymptotic Enumeration of RNA Structures with Pseudoknots

Abstract

In this paper we present the asymptotic enumeration of RNA structures with pseudoknots. We develop a general framework for the computation of exponential growth rate and the sub exponential factors for $k$-noncrossing RNA structures. Our results are based on the generating function for the number of $k$-noncrossing RNA pseudoknot structures, ${\sf S}_k(n)$, derived in \cite{Reidys:07pseu}, where $k-1$ denotes the maximal size of sets of mutually intersecting bonds. We prove a functional equation for the generating function $\sum_{n\ge 0}{\sf S}_k(n)z^n$ and obtain for $k=2$ and $k=3$ the analytic continuation and singular expansions, respectively. It is implicit in our results that for arbitrary $k$ singular expansions exist and via transfer theorems of analytic combinatorics we obtain asymptotic expression for the coefficients. We explicitly derive the asymptotic expressions for 2- and 3-noncrossing RNA structures. Our main result is the derivation of the formula ${\sf S}_3(n) \sim \frac{10.4724\cdot 4!}{n(n-1)...(n-4)} (\frac{5+\sqrt{21}}{2})^n$.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Emma Y. Jin, Christian M. Reidys. 2007-06-21. Asymptotic Enumeration of RNA Structures with Pseudoknots. https://arxiv.org/abs/0706.3137

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

RNA Design via Conditioned Flow Matching and Finite-Policy Reinforcement Learning

RNA design aims to identify sequences that fold into specified secondary structures. Existing methods formulate the task as target-specific search or conditional generation. However, natural RNA evolution proceeds through sequence variation and selection, with compensatory substitutions, whereas these methods do not explicitly model this process. To address this limitation, we propose a two-stage framework comprising RNA Inverse-Folding Flow (RNA-IFlow) and RNA-IFlow-RL. RNA-IFlow uses structure-conditioned Dirichlet Flow Matching to model coordinated variation across the sequence, while RNA-IFlow-RL maps the learned flow to a pairing-preserving finite policy and refines it with thermodynamic feedback. Our framework achieves leading performance on multiple benchmarks, reaching 85.19% Pass@1 on Rfam-27. Further analyses reveal thermodynamic gains, policy dynamics, and robustness across settings. Our work couples coordinated variation with thermodynamic selection, offering a novel paradigm for RNA design.

q-bio.BM↗

Leveraging secondary-structure information for accurate nucleic acid structure prediction with OFoldNA

Recent advances in biomolecular structure prediction have enabled accurate modelling of increasingly complex molecular systems. However, nucleic acid structure prediction remains challenging because of conformational flexibility and the limited availability of high-quality 3D structural data. Secondary structure (SS) provides a more readily available layer of structural information that captures base-pairing relationships and folding topology. Here we present OFoldNA, an all-atom diffusion model that incorporates SS information into nucleic acid folding and protein--nucleic acid co-folding. Without external SS information, OFoldNA achieved leading performance on FoldBench for both nucleic acid monomer folding and protein--nucleic acid co-folding, with particularly strong performance on DNA monomers and protein--DNA interfaces involving longer nucleic acid chains. When accurate base-pairing information was provided, OFoldNA-SS2TS further improved both folding and co-folding accuracy, while partial SS information also yielded consistent gains. The same auxiliary branch can also be used for RNA SS prediction as OFoldNA-SS, which achieved the best out-of-distribution performance on CHANRG. Together, these results show that intermediate structural information such as nucleic acid SS can be leveraged to improve all-atom 3D modelling, providing a general direction for incorporating complementary structural modalities into molecular structure prediction and design.

q-bio.BM↗

How 'Foundational' Are Current Molecular Foundation Models?

Large-scale models have permeated the molecular sciences, yet what makes a model 'foundational' in this domain remains poorly defined. This paper proposes three testable criteria for assessing the foundational nature of molecular models: (i) generality across molecular entities, properties, and tasks; (ii) transferability to new applications with no or minimal task-specific retraining; and (iii) generalization beyond the training distribution. Applying these criteria to the state of the art reveals promising progress, particularly visible in biomolecular structure prediction and machine-learned interatomic potentials, although none of the approaches examined fully satisfies all three. Success is concentrated in domains where target properties are consistently defined and training data are abundant, with low label noise relative to physically meaningful variation. More broadly, progress in the molecular sciences appears to depend less on model scale alone than on the quality and structure of available data, as well as the incorporation of prior knowledge into models, prediction tasks, or downstream applications. This work shifts the notion of a molecular foundation model from a descriptive label to a testable hypothesis, offering a framework for assessing current models and guiding future developments.

q-bio.BM↗