arXiv · 2412.20218
YAD: Leveraging T5 for Improved Automatic Diacritization of Yorùbá Text
Abstract
In this work, we present Yorùbá automatic diacritization (YAD) benchmark dataset for evaluating Yorùbá diacritization systems. In addition, we pre-train text-to-text transformer, T5 model for Yorùbá and showed that this model outperform several multilingually trained T5 models. Lastly, we showed that more data and larger models are better at diacritization for Yorùbá
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Akindele Michael Olawole, Jesujoba O. Alabi, Aderonke Busayo Sakpere, David I. Adelani. 2024-12-28. YAD: Leveraging T5 for Improved Automatic Diacritization of Yorùbá Text. https://arxiv.org/abs/2412.20218
Cite the original work for its findings. Save a collection to share your selection of sources.