arXiv · 2608.28916
VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
Abstract
Automatic speech recognition is usually evaluated with word error rate (WER), although voice workflows often require exact written values. VoiceCodeBench measures whether transcripts preserve identifiers, paths, commands, and other structured tokens needed by downstream software. It contains 300 human-recorded English workplace segments (5.59 hours, 85 speakers) and 1,482 audited entities across 26 types and eight domains. Under a raw-audio-only protocol, we evaluate 19 batch and streaming systems using WER, Canonical Token/Entity Match (CTEM), and strict segment-level Task Success Rate (TSR). Across systems, WER has little rank agreement with CTEM (Spearman $ρ=-0.28$) or TSR ($ρ=-0.22$). The best CTEM and TSR are 91.8% and 68.7%. Even the strongest system therefore leaves nearly one-third of recordings with an unrecovered critical value. Symbol-, separator-, and boundary-sensitive entities account for most errors.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tyler Baumgartner, Brandon Tai, Lisa Kaelin-Martin, Candice Fan, Luc Debaupte, Bill Wang, Yi Zhong. 2026-09-12. VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition. https://arxiv.org/abs/2608.28916
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.