arXiv · 2305.17690
HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language
Abstract
This paper presents HaVQA, the first multimodal dataset for visual question-answering (VQA) tasks in the Hausa language. The dataset was created by manually translating 6,022 English question-answer pairs, which are associated with 1,555 unique images from the Visual Genome dataset. As a result, the dataset provides 12,044 gold standard English-Hausa parallel sentences that were translated in a fashion that guarantees their semantic match with the corresponding visual information. We conducted several baseline experiments on the dataset, including visual question answering, visual question elicitation, text-only and multimodal machine translation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shantipriya Parida, Idris Abdulmumin, Shamsuddeen Hassan Muhammad, Aneesh Bose, Guneet Singh Kohli, Ibrahim Said Ahmad, Ketan Kotwal, Sayan Deb Sarkar, Ondřej Bojar, Habeebah Adamu Kakudi. 2023-05-28. HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language. https://arxiv.org/abs/2305.17690
Cite the original work for its findings. Save a collection to share your selection of sources.