arXiv · 2609.39030
SURE-EVAL: A Systematic and Unified Agentic Framework for Reproducible Evaluation
Abstract
Audio and speech models are released rapidly, but reported scores often conflate model capability with deployment and evaluation choices. The same checkpoint can produce different predictions under different runtimes, hardware, decoding settings, or fallback policies. Even fixed predictions can receive different scores under different normalization and metric implementations. Existing speech benchmarks standardize selected datasets or scoring procedures, but rarely connect heterogeneous model onboarding, controlled inference, and versioned scoring in one executable workflow. We introduce SURE-EVAL, a Systematic and Unified Agentic framework for Reproducible Evaluation of audio and speech systems. A Tool Agent Workflow converts model releases into isolated, verified callable tools. A Main Agent Workflow commits task-specific inference and scoring protocols, then delegates all score-bearing operations to versioned deterministic programs. Each result retains its runtime, protocol, pipeline nodes, predictions, and audit artifacts. Across 18 public releases covering automatic speech recognition, text-to-speech, voice conversion, speaker diarization, speaker-attributed recognition, and multi-task audio understanding, a Codex-only baseline completes 12 models in one shot, while the same agent with the SURE-EVAL Tool Agent Workflow completes all 18. We also conduct unified evaluations over seven ASR test conditions and two TTS subsets. A protocol analysis of three TTS systems finds absolute differences of 0.02-0.52 points between paper-reported and unified results, with the direction varying by model and language. These results show that reproducible evaluation requires controlling both model execution and output scoring.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jing Peng, Junhao Du, Yixuan Wang, Bowen Wang, Hanqi Li, Chaolei Liu, Weihan Chen, Haohui Xie, Ruichen Sun, Chenghao Wang, Wen Wen, Guanyu Chen, Xiaoyu Gu, Haoyu Li, Yiwei Guo, Bohan Li, Tao Liu, Yucheng Wang, Yu Xi, Yihua Zhou, Qiang Zhou, Feng Lu, Shuai Wang, Kai Yu. 2026-09-30. SURE-EVAL: A Systematic and Unified Agentic Framework for Reproducible Evaluation. https://arxiv.org/abs/2609.39030
Cite the original work for its findings. Save a collection to share your selection of sources.