Search arXivSearch

arXiv · 2608.29751

Large-Scale Qualitative Research with AI: Infrastructure, Management and Operation of the Socioscope Data Pipeline

Abstract

The Socioscope project is a pioneering effort in Large-Scale Qualitative Research (LSQR) collecting comparable, open-ended, multimedia field data on hundreds of cases and using AI to make the material analysable at scale. The domain studied is the food system. The entities documented are the organisations that act in it: farms, processors, distributors, retailers, restaurants; and, at meso level, the actors that shape their environment, such as municipalities, government programmes, banks, NGOs and universities. This paper provides the technical reference for how the resulting data Corpus was built and managed to enable AI-augmented analysis. It describes the data pipeline end to end: the systemic sampling frame; the transaction grid used to capture each initiative's relations within the food system; the social contract that rewards participating interviewees, aiming to sustain access; the operational chain from scouting to interviews, including their uploading, transcription, translation, quality control and curation; the provenance rules (originals are immutable, every transformation is logged); and the installation of equipment, personnel and processes, including ethics and GDPR compliance. In its first phase (2023-2026) the pipeline produced 686 documented cases from 31 countries: some 1,430 hours of recordings, about 450,000 speech turns, and 12.6 million words of transcript. We report costs, metrics, lessons learned and limitations, so that other teams can reuse, adapt, and improve the Socioscope methodology.

Explore related subjects

Keep this discovery

BibTeXRIS

Saadi Lahlou, Juan Pablo Caicedo, Shriya Sekhsaria, Valentine Fournand, Paulius Yamin, Helga Nowotny. 2026-08-30. Large-Scale Qualitative Research with AI: Infrastructure, Management and Operation of the Socioscope Data Pipeline. https://arxiv.org/abs/2608.29751

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related discoveries

From Design Principles to Prototype: A Game for Students with ADHD and Learning Disabilities Transitioning to Post-Secondary Education

Students with Attention Deficit Hyperactivity Disorder (ADHD) and Learning Disabilities (LD) can face significant academic, social, and organizational challenges when transitioning to post-secondary education. This paper presents a literature-informed serious game prototype designed to support this transition. We synthesize prior work into design considerations for students with ADHD and LD and show how these considerations are instantiated in a story-driven game.

cs.MM

BIRDS: Characterizing and Understanding Biodiversity Impact of Large Language Model Serving

Large language model (LLM) serving creates environmental impacts beyond carbon and water, including ecosystem damage through biodiversity-related pathways. We present BIRDS, a framework for Biodiversity Impact of Request-Driven LLM Serving. BIRDS defines request-level functional units, quantifies operational and embodied biodiversity impact, and introduces Quality-Normalized Biodiversity Impact (QNBI) to jointly analyze ecological impact and response quality. Across diverse workloads, models, GPUs, and regions, BIRDS reveals that biodiversity impact accumulates at scale and exposes quality-aware serving tradeoffs. The code is available at https://github.com/TianyaoShi/BIRDS.

q-bio.OT

Experts Disagree on How to Fight AI Disinformation, but Agree That Health and Politics Need Different Solutions

When 54 international experts assessed AI-generated disinformation threats, they revealed a surprising pattern: while video deepfakes received the highest average threat ratings in the political domain (M = 6.31/7), the pattern differed in the health domain, where AI-generated text received the highest average rating (M = 5.80). Experts also diverge on what to do: government regulation drew both the most "most effective" (30%) and the most "least effective" (15%) votes, though rating distributions were contested rather than polarized, indicating disagreement over priorities rather than over efficacy. These findings offer an initial expert map of an AI-disinformation landscape that is still rapidly forming.

cs.CY