arXiv · 1503.08167
Normalization of Non-Standard Words in Croatian Texts
Abstract
This paper presents text normalization which is an integral part of any text-to-speech synthesis system. Text normalization is a set of methods with a task to write non-standard words, like numbers, dates, times, abbreviations, acronyms and the most common symbols, in their full expanded form are presented. The whole taxonomy for classification of non-standard words in Croatian language together with rule-based normalization methods combined with a lookup dictionary are proposed. Achieved token rate for normalization of Croatian texts is 95%, where 80% of expanded words are in correct morphological form.
Explore related subjects
Keep this discovery
Slobodan Beliga, Miran Pobar, Sanda Martinčić-Ipšić. 2015-03-27. Normalization of Non-Standard Words in Croatian Texts. https://arxiv.org/abs/1503.08167
Cite the original work for its findings. Save a collection to share your selection of sources.