Search arXiv⌕ Search

arXiv · 2610.07607

Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution

Abstract

Model-guided directed evolution seeks to identify high-fitness protein variants under limited oracle budgets. Protein language models (PLMs) provide rich representations for this task, but task-agnostic zero-shot scores can be misaligned with a target assay, while supervised search in high-dimensional embedding spaces can make surrogate modeling and uncertainty estimation sample-inefficient. We propose the Linear Fitness Subspace (LFS) hypothesis: within mutation-induced residue-level representation changes, a compact, assay-specific set of directions makes fitness variation linearly accessible from few labeled variants. This is a local, supervision-recoverable statement rather than a claim that protein fitness landscapes or global PLM geometry are universally linear. Building on this observation, we introduce Subspace-Guided Evolutionary Search (SGES), which estimates an LFS from a small initial sample and performs surrogate modeling, uncertainty estimation, and acquisition in the learned subspace. Across 10 core ProteinGym assays, 87 extended static-validation assays, and an 18-assay budgeted-search evaluation, SGES improves fitness prediction and search efficiency over zero-shot PLMs and recent ML-guided protein optimization baselines. Controlled comparisons with PCA, random projections, label-shuffled PLS, classical mutation features, and acquisition ablations further isolate the benefit of a fitness-aligned site-delta coordinate.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

SiYuan Ma, Canran Xiao, Zikai Xiao, Albert Gao, Liang He, Xuan-Yu Wang, Shuying Cao, Xiaojun Jia. 2026-10-06. Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution. https://arxiv.org/abs/2610.07607

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

How causal analysis can reveal autonomy in models of biological systems

Standard techniques for studying biological systems largely focus on their dynamical, or, more recently, their informational properties, usually taking either a reductionist or holistic perspective. Yet, studying only individual system elements or the dynamics of the system as a whole disregards the organisational structure of the system - whether there are subsets of elements with joint causes or effects, and whether the system is strongly integrated or composed of several loosely interacting components. Integrated information theory (IIT), offers a theoretical framework to (1) investigate the compositional cause-effect structure of a system, and to (2) identify causal borders of highly integrated elements comprising local maxima of intrinsic cause-effect power. Here we apply this comprehensive causal analysis to a Boolean network model of the fission yeast (Schizosaccharomyces pombe) cell-cycle. We demonstrate that this biological model features a non-trivial causal architecture, whose discovery may provide insights about the real cell cycle that could not be gained from holistic or reductionist approaches. We also show how some specific properties of this underlying causal architecture relate to the biological notion of autonomy. Ultimately, we suggest that analysing the causal organisation of a system, including key features like intrinsic control and stable causal borders, should prove relevant for distinguishing life from non-life, and thus could also illuminate the origin of life problem.

q-bio.QM↗

easyplater: The easy way to generate microplate designs deconvolved from multivariate clinical data

Microplate-based omic studies of large clinical cohorts can accelerate biomedical research, but experimental power and veracity are compromised when plate positional effects confound clinical variables of interest. Plate designs must therefore deconvolve positional and biological variation, but existing computational approaches still rely on manual intervention to ensure adherence to spatial constraints. Here, we present three complementary advances that reduce researcher effort. First, we propose a weighted, multivariate plate design score comprising a novel metric of spatial autocorrelation that rewards global separation of similar samples, and a penalty for local, variable-wise homogeneous regions. Second, we use a network-based approach to identify clinically similar samples, generate random layouts under the constraint that similar samples are allocated to distal wells, and select the top-scoring candidate as an initial layout. Lastly, we use an efficient sample-swapping search to improve this layout. We implemented this method in easyplater, an R package for generating 96-well plate designs that takes clinical data and variable weights as input and outputs the top-scoring layout in CSV, XLSX or HTML format. Overall, easyplater substantially reduces user intervention, outperforms existing methods, and facilitates robust plate design for large, plate-based omic studies.

q-bio.QM↗

APOD: reasoning-guided agentic population ordinary differential equation discovery for pharmacological digital twins

Establishing ordinary differential equations (ODEs) describing population data is a fundamental part of mathematical modeling in pharmacology, crucial to developing digital twins. However, doing so from sparse, noisy data is a slow, expert-driven task. Existing automated methods either search a restricted model space or ignore population inter-individual variability. Here we introduce APOD (Agentic Population ODE Discovery), a language-model agent that iteratively reasons over biological knowledge and fit diagnostics in an open-ended search space to discover a population digital twin (PDT), i.e., a shared ODE system with between-subject variability. On synthetic pharmacokinetic and tumor-dynamics benchmarks, APOD recovered ground-truth structures in 94-100\% of runs, 12-fold faster in median than an established library-based search. On real cohorts it converged to valid structures, and proposed a PDT of radioligand-therapy-induced platelet dynamics that predicts thrombocytopenia from first-cycle data and simulates alternative dosing schedules that lower the predicted risk of toxicity.

q-bio.QM↗