arXiv · 2609.32527
AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses
Abstract
Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolute metrics cannot separate target-specific predictions from a shared background response. Second, common metrics remain high under gene shuffling, so gene-level accuracy is never verified. Third, a score at one training size says nothing about coverage, which depends on representation-space proximity and response-constraining power. We propose AmbiModBench, a specificity-aware, gene-resolved and coverage-aware benchmark. It pairs every score with a training-mean reference fitted on the same split, screens each readout by gene-coordinate permutation, and links embedding distance to response variation. Across K562, RPE1 and Norman, strong absolute scores largely reflect shared background rather than target-specific learning. Widely used readouts track response magnitude distributions rather than the affected genes. Detectable gain follows representation-space coverage rather than training-set size. Nonetheless, on RPE1 the protocol yields a reproducible target-specific gain across five additional splits and three gene selections, which absolute scores alone cannot distinguish from shared background.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sikai Huang, Zhiwen Yang, Kai Yu, Jiayuan Chen, Stan Z. Li. 2026-09-26. AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses. https://arxiv.org/abs/2609.32527
Cite the original work for its findings. Save a collection to share your selection of sources.