Search arXiv⌕ Search

arXiv · 2609.36834

WenetSpeech-Min: A Large-Scale Minnan Speech Corpus with Dual Transcriptions for Dialectal Speech Processing

Abstract

Progress in dialectal speech technology is hindered by the scarcity of large-scale, real-world corpora. For Minnan speech, existing resources remain limited, and few provide paired Minnan and Mandarin transcripts at scale. To address these gaps, we introduce WenetSpeech-Min, an open-source corpus comprising around 10,000 hours of Minnan speech collected from diverse online media, with paired Minnan and Mandarin transcripts for every utterance. We further establish an automatic speech recognition (ASR) benchmark covering both Minnan and Mandarin transcripts and a text-to-speech synthesis (TTS) benchmark using Minnan transcripts, with manually verified evaluation sets for both tasks. To assess the effectiveness of the corpus, we train ASR and TTS models on WenetSpeech-Min and compare them with representative systems on the proposed benchmarks. The resulting models outperform the evaluated open-source models on most metrics and achieve competitive performance against commercial systems. We will release the corpus, benchmarks, and models to facilitate reproducible research on Minnan speech technology.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Haoyu Zhang, Chunjiang He, Hongtao Li, Zeyu Zhu, Qituan Shangguan, Chengyou Wang, Jingbin Hu, Ziyu Zhang, Bingshen Mu, Yanbo Wang, Shuai Wang, Jinhui Ye, Chengdong Liang, Binbin Zhang, Pengcheng Zhu, Chuang Ding, Qianze Feng, Qingyang Hong, Liumeng Xue, Lei Xie. 2026-09-29. WenetSpeech-Min: A Large-Scale Minnan Speech Corpus with Dual Transcriptions for Dialectal Speech Processing. https://arxiv.org/abs/2609.36834

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios

Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora for training. Synthetic visual data have been shown to be an effective augmentation strategy for addressing AV data scarcity. However, a more challenging scenario arises for languages such as Catalan, where no real audiovisual data are available for training. In this study, we investigate whether AVSR can be bootstrapped in such a zero-AV-resource setting, using synthetic visual data as the sole source of visual supervision. We synthesize over 700 hours of talking-head video and fine-tune a pre-trained AV-HuBERT model. On a manually annotated Catalan benchmark, our model achieves near state-of-the-art (SOTA) performance with much fewer parameters and training data than SOTA ASR systems such as Whisper-large-v3, outperforms an identically trained audio-only baseline, and preserves multimodal advantages under acoustic degradation. Scalable synthetic video thus offers a viable substitute for real recordings in zero-AV-resource AVSR.

eess.AS↗

Teacher-Free Self-Distilled Consistency Trajectory Learning for Fast Speech Enhancement

Consistency trajectory models offer a route to fast, high- quality speech enhancement, collapsing the many reverse steps of diffusion-based enhancers into a handful. When instantiated on a Schrödinger bridge (SB), which pins the generative process to fixed clean and noisy endpoints, exist- ing consistency-trajectory enhancers (SBCTMs) still require a pretrained teacher to supply trajectory supervision, which raises training cost and ties the final quality to that of the teacher. We propose a teacher-free, self-distilled consistency- trajectory framework that removes the external teacher result- ing in a 5X reduction in per epoch training time. Our model is trained with a three-stage curriculum of clean speech pre- diction, a self-distilled shortcut objective, and perceptual fine-tuning with a multi-resolution short-time Fourier trans- form (MR-STFT) loss. Using the same NCSN++ backbone as SBCTM, our model attains a wide-band PESQ of 3.01, ES- TOI 0.87, and SI-SDR 19.07 dB on VoiceBank+DEMAND compared to 3.57, 0.87 and 12.8 dB for the teacher based model. Further, we find that a geometric schedule at low reverse step count maximizes perceptual quality, while a higher-step uniform schedule favors signal fidelity.

eess.AS↗

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1$\times$ with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10$\times$ faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.

eess.AS↗