Search arXiv⌕ Search

arXiv subjects

Ming-Hsiu Wu

Publications and source records attributed to Ming-Hsiu Wu.

2 recordsLinked to original sources

Assay-Aware BindingDB: Curating Experimental Context for Binding Affinity Prediction

Protein--ligand binding affinity prediction is fundamental to computational drug discovery, yet modern AI-driven models are limited by pervasive heterogeneity in their training data: bioactivity values are aggregated across diverse assay types and experimental conditions without accounting for protocol-level differences, introducing systematic noise. Existing harmonization approaches either discard assay-level metadata or collapse it into coarse categorical distinctions, leaving rich contextual signal unused. We address this gap with two contributions. First, we introduce Assay-Aware BindingDB, augmenting 74,425 BindingDB protein--ligand pairs across four assay types (ITC, SPR, RBA, and FPA) with structured metadata extracted from primary literature using a two-stage agentic framework that separates evidence extraction from ontology-conditioned JSON synthesis. Against domain-expert references, the framework achieves presence F1 $\geq 0.947$ and semantic content accuracy $\geq 0.913$ across all assay types. Second, we develop a context-conditioned affinity model that injects an assay-context embedding into the Boltz-2 affinity module. On a paper-level held-out split, the model reduces combined-regime MSE from $1.27$ to $1.19$ and increases Pearson correlation from $0.64$ to $0.67$, with significant gains for SPR and FPA. ITC, the only label- and immobilization-free assay considered, shows no improvement, consistent with its design removing the protocol artifacts the metadata captures. These results support the hypothesis that systematically curated assay metadata provides informative signal for affinity prediction.

q-bio.BM↗

Towards Precision Protein-Ligand Affinity Prediction Benchmark: A Complete and Modification-Aware DAVIS Dataset

Advancements in AI for science unlocks capabilities for critical drug discovery tasks such as protein-ligand binding affinity prediction. However, current models overfit to existing oversimplified datasets that does not represent naturally occurring and biologically relevant proteins with modifications. In this work, we curate a complete and modification-aware version of the widely used DAVIS dataset by incorporating 4,032 kinase-ligand pairs involving substitutions, insertions, deletions, and phosphorylation events. This enriched dataset enables benchmarking of predictive models under biologically realistic conditions. Based on this new dataset, we propose three benchmark settings-Augmented Dataset Prediction, Wild-Type to Modification Generalization, and Few-Shot Modification Generalization-designed to assess model robustness in the presence of protein modifications. Through extensive evaluation of both docking-free and docking-based methods, we find that docking-based model generalize better in zero-shot settings. In contrast, docking-free models tend to overfit to wild-type proteins and struggle with unseen modifications but show notable improvement when fine-tuned on a small set of modified examples. We anticipate that the curated dataset and benchmarks offer a valuable foundation for developing models that better generalize to protein modifications, ultimately advancing precision medicine in drug discovery. The benchmark is available at: https://github.com/ZhiGroup/DAVIS-complete

cs.LG↗