Search arXiv⌕ Search

arXiv · 2610.02924

Evaluator-in-the-Loop Monte Carlo Tree Search via LLM Agents for Motif Scaffolding in Protein Design

Abstract

Motif-scaffolding systems commonly follow a generate-then-filter paradigm, in which candidate proteins are generated independently and structural evaluation is used primarily for terminal screening or ranking. This paradigm underuses evaluation: failed predictions contain state-specific evidence about whether a design requires repair of motif geometry, global foldability, or other structural constraints. We introduce \textbf{ELMS} (Evidence-based LLM-guided Monte Carlo Search), an evaluator-in-the-loop search framework for motif scaffolding that turns such evaluator feedback into targeted design actions. Effective reuse of structural feedback is nontrivial because different scaffold states exhibit different failure modes, and repeatedly refining a single trajectory can prematurely commit computation to an unproductive region of sequence space. ELMS therefore retains evaluated scaffolds as persistent search states: a Critic Agent diagnoses state-local structural failures, a Policy Agent selects targeted operators with execution parameters, motif-locked operators realize legal sequence modifications, and MCTS determines which historical states should receive further design effort. Under the standard GeomMotif protocol (100 candidates per task), ELMS achieves Successful rates of 86.41\% on single-motif tasks and 84.57\% on paired-motif tasks, exceeding the strongest prior baseline by 19.3 and 21.9 percentage points, respectively. On MotifBench, under a matched 100-candidate search budget, it solves 26.7 of 30 tasks on average (88.89\% Task Success), compared with 16.0 tasks (53.33\%) for the strongest baseline. These results establish ELMS as an effective approach for converting structural evaluation from a terminal filter into actionable guidance for iterative motif scaffolding.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Haotian Hu, Oguzhan Gungordo, Siheng Xiong, Faramarz Fekri. 2026-10-02. Evaluator-in-the-Loop Monte Carlo Tree Search via LLM Agents for Motif Scaffolding in Protein Design. https://arxiv.org/abs/2610.02924

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Active Learning for Communication Structure Optimization in LLM-Based Multi-Agent Systems

Optimizing the communication structure of large language model based multi-agent systems (LLM-MAS) has been shown to improve downstream performance and reduce token usage. Existing methods typically rely on randomly sampled training tasks. However, tasks may differ substantially in difficulty and domain, and thus they are not equally informative for updating communication structure, making optimization often unstable and highly sensitive to the particular training set. To actively identify the most valuable tasks for communication-structure optimization, we propose an ensemble-based information-theoretic task selection framework. The proposed method estimates task informativeness by how much a candidate task changes the distribution over graph parameters, using ensemble Kalman inversion as an efficient and derivative-free approximation of the corresponding Bayesian update. The resulting estimator is especially suitable for black-box and noisy multi-agent systems. To enhance scalability, we construct a compact candidate pool through embedding-based representative selection and combine the informative selection with surrogate modeling and batch Thompson sampling. We validate the proposed framework across both benign and adversarial settings and multiple task formats. It consistently outperforms random training, demonstrating more effective task selection and greater overall cost efficiency.

cs.MA↗

You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

Ensuring aligned agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. We show that a single aligned agent can propagate cooperative behaviors to unmodified agents purely through natural-language interaction, a phenomenon we term Alignment Propagation. We study this in the Red-Black Game, a team-based iterated Prisoner's Dilemma in which teammates deliberate and vote to determine their team's collective action. By distilling the cooperative reasoning and persuasive dialogues of a teacher model into Qwen3-14B, we obtain a seed agent that, when placed among four unmodified teammates, more than doubles the cooperation rate from 24.8% to 62.2%, outperforming the teacher model and a vanilla Gemini-3.1-Pro. Remarkably, a seed trained exclusively on the Red-Black Game transfers zero-shot to Sugarscape, a spatially grounded survival simulation with pairwise trading, achieving a 91.5% trade success rate versus a 21.6% baseline. Our results reframe multi-agent alignment from an exhaustive per-agent training problem to a scalable social capability that can be engineered through strategic seed placement.

cs.MA↗

Cache Merging as a Convergent Replicated State for Multi-Agent Latent Reasoning

Multi-agent latent reasoning composes the KV-cache contributions of several agents into one context for a final agent. Prior work (Agent Primitives) does this by concatenating per-agent caches with RoPE re-encoding, a construction we name BagMerge. BagMerge is non-commutative, and the best input ordering is not predictable a priori: it shifts with deployment regime, latent-step budget, model scale, and model family. We make this cache exchange a convergent replicated state. CanonicalMerge fixes the layout by a content-determined ordering (mean K-norm at a middle layer), making the merged cache byte-identical under any input permutation, verified on synthetic tensors (N <= 5) and bit-for-bit on real Qwen3-1.7B and Qwen3-4B KV state. We then separate state from layout: the durable object is a set of content-addressed latent fragments merged by set union, a state-based CvRDT, and CanonicalMerge is its deterministic render, so every accuracy number is inherited and re-delivered duplicates are absorbed. On a partitioned-reasoning benchmark CanonicalMerge matches the best BagMerge ordering without knowing which it is (within 4 points in all 12 cells at 1.7B; same picture at 4B), and on Llama-3.1-8B, where the ordering gap grows to 19 points, it again tracks the best ordering. On HotpotQA (n = 200) and MuSiQue it is the best cache-level method, at a 7-point F1 cost against shipping the text on HotpotQA and at parity on MuSiQue, while the output-fusion baseline PackLLM trails by 45 points. At k > 2 we delimit the approach: cache merge transports latent traces but does not by itself compose them.

cs.MA↗