arXiv · 2609.29371
BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech
Abstract
This paper presents BanglaTurn, a corpus for end-of-turn detection in Bangla conversational speech, and a model trained on it. The corpus holds 35,374 samples of 3 to 15 s of podcast speech, labelled for turn state by combining speaker diarization with an LLM pass, with every label then checked by a human annotator. The model pairs a Whisper encoder with task-specific classification heads. On a class-balanced test set drawn from a held-out podcast, it reaches 84.33% accuracy (95% CI 80.3 to 88.1) against 69.28% for the Smart-Turn v3 baseline, and lowers the false negative rate from 51.57% to 7.55% at the cost of a higher false positive rate. We report what encoder layer fine-tuning, multi-scale pooling and INT8 quantization each contribute, and latency stays within 165 to 191 ms end to end on CPU.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mizbaul Haque Maruf. 2026-09-24. BanglaTurn: A Benchmark and Whisper-Based Model for End-of-Turn Detection in Bangla Speech. https://arxiv.org/abs/2609.29371
Cite the original work for its findings. Save a collection to share your selection of sources.