VibeProteinBench: An Evaluation Benchmark for Language-interfaced Vibe Protein Design
Protein design aims to compose amino-acid sequences that fold into stable three-dimensional structures while satisfying targeted functional properties. The field is increasingly shifting toward vibe protein design, where a single model is expected to generate novel sequences, engineer existing proteins, and reason about protein characteristics through flexible natural-language constraints. However, existing benchmarks typically evaluate isolated aspects of protein design or assume predefined structured inputs, making them ill-suited to assess this broad, open-ended setting. To address this gap, we present Vibe Protein Design Benchmark (VibeProteinBench), a language-interfaced benchmark that evaluates whether a single model can operate across three complementary stages of a computational protein-design workflow: recognition, engineering, and generation. Each stage is grounded in expert-curated mechanistic rationales and multi-faceted in silico validation, to computationally verify whether model outputs are biologically plausible. Across diverse general-purpose and domain-specialized LLMs, evaluated with and without tool access, no model performs strongly across all three stages at once. Beyond measuring end-task success, VibeProteinBench exposes where current systems break: we identify recurring failure patterns and characterize how tool-augmented agents select tool arguments. Together, these findings show that vibe protein design remains a substantial open challenge.