arXiv · 2609.25496
Towards participatory speech dataset curation: A queer case study and conceptual framework
Abstract
In this paper, we motivate the need for a participatory speech dataset creation framework through a case study of the LGBTQIA+, or queer, community - a community with documented concerns about AI and reported harms, including attempts to develop 'gaydar' technologies that purportedly identify individuals as queer. We review common speech data collection practices, why these methods may be unsuitable for engaging with queer speakers, and discuss previous efforts in participatory AI with queer community engagement, as well as participatory endeavours specific to speech data collection for other marginalized communities. From this review, we develop a conceptual framework for participatory speech data curation by, for, and with marginalized communities drawing on insights from co-design and knowledge sharing. We propose a framework comprising overlapping and two-way processes of defining a community, project formulation, modes of participation, and personal autonomy.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Brooklyn Sheppard, Anaelia Ovalle, Adina Williams, Levent Sagun. 2026-09-21. Towards participatory speech dataset curation: A queer case study and conceptual framework. https://arxiv.org/abs/2609.25496
Cite the original work for its findings. Save a collection to share your selection of sources.