arXiv · 2609.31967
IndicFDB: Benchmarking Full-Duplex Voice Agents across Indian Languages
Abstract
Full-duplex voice agents must handle pauses, take turns, backchannel, and respond to user interruptions in real time. Full-Duplex-Bench evaluates these behaviors, but its English-only corpus and reliance on word-timestamped ASR and an English-prompted LLM judge make it difficult to extend to Indian languages. We introduce IndicFDB, which extends it to ten languages spoken in India with 12,350 samples, nearly 17 times as many as the original. We address three challenges: finding conversational events in multilingual speech, evaluating their timing without reliable word-level alignment, and judging responses across languages. We mine pause handling, turn taking, and backchanneling samples from roughly 50,000 hours of channel-separated conversations using voice activity detection (VAD), and construct human-validated synthetic user interruption samples. Language-independent VAD heuristics evaluate timing, while an open-weight transcription and translation pipeline converts responses to English for LLM ratings of relevance and quality. Across seven voice agents, commercial APIs show unexpectedly consistent behavior across languages but are either fast or robust to pauses, never both, while monolingual open full-duplex models expose further tradeoffs among backchanneling, response quality, and latency.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Rajarshi Roy, Shobhit Banga, Jonathan Raiman, Supriya Paul, Bhaskar Singh, Manmeet Kaur, Sagar Jain, Hanuman Sidh, Pranav Sharma, Aditya Singh, Aaditya Pareek, Manas Dhir, Adi Margolin, Niket Agarwal, Bryan Catanzaro. 2026-09-25. IndicFDB: Benchmarking Full-Duplex Voice Agents across Indian Languages. https://arxiv.org/abs/2609.31967
Cite the original work for its findings. Save a collection to share your selection of sources.