arXiv · 1911.07555
Short Text Language Identification for Under Resourced Languages
Abstract
The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text language identification (LID) useful for under resourced languages. The algorithm is evaluated on short pieces of text for the 11 official South African languages some of which are similar languages. The algorithm is compared to recent approaches using test sets from previous works on South African languages as well as the Discriminating between Similar Languages (DSL) shared tasks' datasets. Remaining research opportunities and pressing concerns in evaluating and comparing LID approaches are also discussed.
Explore related subjects
Keep this discovery
Bernardt Duvenhage. 2019-11-18. Short Text Language Identification for Under Resourced Languages. https://arxiv.org/abs/1911.07555
Cite the original work for its findings. Save a collection to share your selection of sources.