Search arXiv⌕ Search

arXiv subjects

Abdurrahman Juma

Publications and source records attributed to Abdurrahman Juma.

2 recordsLinked to original sources

CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR

We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed YouTube speech spanning 21 micro-dialects across Egypt, Jordan, Lebanon, Mauritania, Palestine, and Saudi Arabia. Each segment is annotated with one or more operational micro-dialect labels, code-switching information, and utterance-level perceived gender, enabling analysis of sub-country variation that is obscured by conventional country-level labels. We describe the collection and annotation process, motivate the micro-dialect inventory linguistically, and benchmark four multilingual ASR systems in zero-shot and adapted settings. The strongest zero-shot system obtains 43.47% aggregate WER, with particularly high error rates on Mauritanian and Lebanese varieties; adaptation on Cardamom reduces its WER to 35.21%. Audio-based identification experiments further show that the annotations provide a learnable prediction target, with a dedicated classifier reaching 85.57% accuracy on 21-way micro-dialect identification. Cardamom provides a resource for studying localized dialectal variation and developing Arabic speech systems with broader regional coverage.

cs.CL↗

NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task

We present the findings of the sixth Nuanced Arabic Dialect Identification (NADI 2025) Shared Task, which focused on Arabic speech dialect processing across three subtasks: spoken dialect identification (Subtask 1), speech recognition (Subtask 2), and diacritic restoration for spoken dialects (Subtask 3). A total of 44 teams registered, and during the testing phase, 100 valid submissions were received from eight unique teams. The distribution was as follows: 34 submissions for Subtask 1 "five teamsæ, 47 submissions for Subtask 2 "six teams", and 19 submissions for Subtask 3 "two teams". The best-performing systems achieved 79.8% accuracy on Subtask 1, 35.68/12.20 WER/CER (overall average) on Subtask 2, and 55/13 WER/CER on Subtask 3. These results highlight the ongoing challenges of Arabic dialect speech processing, particularly in dialect identification, recognition, and diacritic restoration. We also summarize the methods adopted by participating teams and briefly outline directions for future editions of NADI.

cs.CL↗