Search arXiv⌕ Search

arXiv · 2506.13395

CBTOPE2: An improved method for predicting of conformational B-cell epitopes in an antigen from its primary sequence

Abstract

In 2009, our group pioneered a novel method CBTOPE for predicting conformational B-cell epitopes in a protein from its amino acid sequence, which received extensive citations from the scientific community. In a recent study, Cia et al. (2023) evaluated the performance of conformational B-cell epitope prediction methods on a well-curated dataset, revealing that most approaches, including CBTOPE, exhibited poor performance. One plausible cause of this diminished performance is that available methods were trained on datasets that are both limited in size and outdated in content. In this study, we present an enhanced version of CBTOPE, trained, tested, and evaluated using the well-curated dataset from Cai et al. (2023). Initially, we developed machine learning-based models using binary profiles, achieving a maximum AUC of 0.58 on the validation dataset. The performance of our method improved significantly from an AUC of 0.58 to 0.63 when incorporating evolutionary information in the form of a Position-Specific Scoring Matrix (PSSM) profile. Furthermore, the performance increased from an AUC of 0.63 to 0.64 when we integrated both the PSSM profile and relative solvent accessibility (RSA). All models were trained, tested, and optimized on the training dataset using five-fold cross-validation. The final performance of our models was assessed using a validation or independent dataset that was not used during hyperparameter optimization. To facilitate scientific community working in the field of subunit vaccine, we develop a standalone software and web server CBTOPE2 (https://webs.iiitd.edu.in/raghava/cbtope2/).

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Anupma Pandey, Megha, Nishant Kumar, Ruchir Sahni, Gajendra P. S. Raghava. 2025-06-16. CBTOPE2: An improved method for predicting of conformational B-cell epitopes in an antigen from its primary sequence. https://arxiv.org/abs/2506.13395

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PocketVE: Stable and Property-Guided Structure-Based Drug Design with Variance-Exploding Diffusion

Protein-conditioned 3D molecule generation is a central challenge in structure-based drug design, requiring a balance between pocket compatibility, molecular properties, and physical geometry. We propose \textbf{PocketVE}, a protein-pocket-conditioned variance-exploding (VE) diffusion framework that couples stable coordinate denoising with inference-time property guidance. Specifically, PocketVE combines an EDM-style training and sampling setup for 3D denoising, classifier-free guidance for multi-property steering without external property classifiers, and adaptive protein perturbation as a training-time pocket regularizer. Evaluated on CrossDocked2020 under the GenBench3D protocol, PocketVE improves Valid$_{3\text{D}}$ from 58.6 to 80.6 and reduces strain energy from 457.4 to 127.9 relative to its TAGMol architectural baseline, while retaining competitive docking and molecular-property scores under moderate guidance. A guidance-scale study shows that moderate guidance gives a favorable balance between target-related objectives and geometric quality, whereas stronger guidance can degrade geometry and distributional fidelity. Pocket-permutation and PoseCheck diagnostics further support pocket-specific spatial compatibility with reduced steric conflicts. Overall, the results suggest that geometric stability and inference-time property guidance should be considered as coupled design objectives.

q-bio.BM↗

In Vivo Length Distributions as Mechanistic Fingerprints of Pathological Protein Aggregation

Modern imaging techniques can resolve individual pathological protein aggregates in postmortem human samples, providing detailed measurements of aggregate size distributions that are inaccessible with conventional bulk approaches. These distributions represent mechanistic fingerprints of the microscopic processes that generated the observed pathology, but extracting this mechanistic information requires a quantitative theoretical framework. Here, we develop the mathematical tools needed to interpret aggregate length distributions in living systems, where aggregate growth competes with active removal. We show that, across a class of models, the length distribution of sufficiently large aggregates approaches a geometric decay. Crucially, the decay rate is determined by the balance between aggregate elongation and removal, providing a direct quantitative readout of these competing processes from a single time point measurement. This enables mechanistic comparisons between healthy and diseased human samples without requiring longitudinal measurements of aggregate dynamics. We further analyse how additional aggregation and removal processes modify the observed length distributions. Together, these results establish the mathematical foundations and tools to use aggregate length distributions as an experimentally accessible route for inferring microscopic aggregation dynamics directly from human tissue.

q-bio.BM↗

Decoding enzyme-substrate interaction topology reveals principles underlying catalytic efficiency and mutational outcomes

The enzyme turnover number (kcat) defines catalytic efficiency and constrains quantitative models of metabolism, yet the molecular determinants governing kcat and its response to mutation remain poorly understood. Measurements are sparse and labor-intensive, and most computational approaches provide numerical predictions without explaining how enzyme-substrate interactions shape catalytic outcomes. A central challenge is therefore to identify the topological principles that determine where mutations act and how their functional outcomes are encoded within the enzyme-substrate interaction network. Here, we show that catalytic efficiency and mutational effects can be interpreted through enzyme-substrate interaction topology. We developed Interkcat, an interpretable bidirectional cross-attention framework that captures reciprocal coordination between protein residues and substrate atoms. Optimized on a unified benchmark, Interkcat achieves state-of-the-art predictive performance (R2 = 0.701). From its learned representations, we derive an Interaction Topology Score (ITS) that identifies sequence regions statistically enriched for mutation-sensitive sites without explicit structural inputs. We further demonstrate that higher-order topological features distinguish opposing mutational outcomes: lethal mutations disrupt coordinated networks, whereas activity-preserving or enhancing mutations retain sparse, globally organized coupling. These findings establish interaction topology as a unifying principle linking enzyme sequence, catalytic efficiency, and evolutionary perturbation.

q-bio.BM↗