Search arXiv⌕ Search

arXiv · 2610.01463

Learning to structure data from user-generated thematic corpora

Abstract

Thematic corpora, such as social media communities, contain unstructured text describing data that could be made structured. These include, for example, personal attributes, behaviors, and experiences mentioned in social media data. Extracting structured data is challenging as relevant attributes are often implicit, domain-dependent, and unknown in advance. We propose a fully automated, iterative framework for discovering and extracting domain-specific attribute schemas without a predefined ontology. Using large language models (LLMs), the framework induces candidate attributes, sequentially consolidates semantically overlapping attributes, and assigns a structural type. These enable creating an ontology and populating it with values from the corpus. The framework also enables the use of smaller LLMs for value extraction with estimable accuracy loss compared to large LLMs. We evaluate the framework on 5 health-related Reddit communities. Discovered attributes achieved 61% agreement with human-identified attributes, close to the 62% agreement between independent annotators. In most cases, the algorithm converges to a stable attribute set in fewer than 10 iterations. Structural type assignment achieves 82% accuracy, and value extraction reaches an F1 score of 0.8 compared to human annotations. Across four LLM families, smaller instruction-tuned models show statistically significant improvements in extraction performance with model scale when evaluated against a high-capacity reference LLM, supporting informed accuracy-cost trade-offs. These results show that attributes comparable to those identified by humans can be discovered automatically, enabling the creation of high-quality structured datasets economically and at scale. By removing the need for predefined ontologies, iterative model-driven schema induction offers a practical and scalable foundation for mining thematic corpora.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Elishay Avram, Oren Glickman, Elad Yom-Tov. 2026-10-01. Learning to structure data from user-generated thematic corpora. https://arxiv.org/abs/2610.01463

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Route What Remains: A Meta-Modal Agent for Missing-Modality Candidate Reranking in Recommender Systems

Missing-modality recommenders usually reconstruct absent representations, although the observed evidence may not determine the missing content. We formulate candidate reranking as budgeted sequential evidence acquisition. A policy queries text, image, and interaction-graph tools, incorporates \texttt{Null} returns into its observation history, and sparsely rescores a retrieved candidate pool. Our \textbf{Meta-Modal Agent} (MMA) uses PPO to optimize terminal NDCG and tool cost without explicit access to the route-availability mask or target identity. When only one evidence route is available, MMA-Auto improves NDCG@10 by $10.0$\% over the strongest completion baseline and by $9.5$\% over a fixed router with the same Llama scorer. It obtains the highest result in all nine reported combinations of dataset and available route against these comparators. MMA-Auto also reduces failed calls by 17.8 percentage points and uses 1.1 fewer turns than the fixed router. On the fixed candidate pools produced by full-catalog retrieval, MMA-Auto improves NDCG@10 by $19.7$\%. These results associate adaptive evidence routing with improved reranking under severe, constructed missingness. The code is available at: https://anonymous.4open.science/r/WSDM2027-MMA-C381.

cs.IR↗

TAGGRAPH: Tag-Augmented Graphs for Graph Retrieval of Agent Persistent Histories

Long-term memory lets LLM agents recall past interactions and remain consistent across sessions, but memory systems are hard to compare because they often vary in representation, indexing, retrieval, and evaluation. We present a controlled evaluation framework based on shared 5W-style conversational memories. Localized graph configurations traverse a common base graph; AdaptiveGraph adds chronological edges and Personalized PageRank diffusion. We also evaluate BM25 over the same extracted notes and OpenClaw as a raw-input external reference. Retrieval rankings vary across memory settings. On LongMemEval-S, AdaptiveGraph is the strongest graph configuration at 0.844 MRR, but BM25 reaches 0.867 and OpenClaw 0.880. On ATANT Core, localized graph traversal outperforms diffusion and BM25, whereas BM25 leads the stress rounds. Reducing LongMemEval-S within the tested range does not reproduce the ATANT diffusion penalty, but the smallest tested store remains larger than ATANT Core, so store size cannot be ruled out. The penalty also persists under a permissive content-match criterion. Vocabulary normalization and extraction quality substantially affect graph retrieval, and missing extraction tags are common among top-five misses. Retrieval strategies should therefore be evaluated jointly with the memory setting and against strong lexical baselines.

cs.IR↗

Exploring Forum Post Retrieval with Generative Modeling

Generative recommendation (GR) has emerged as an alternative to embedding-based retrieval, building on the success of generative models in language and vision. We are exploring GR on Facebook Forum, a standalone application for medium-to-heavy users of Facebook Groups. Because Forum is a new surface, its own interaction data are too sparse to train a GR model from scratch. We address this with transfer along two axes: we train on a broader corpus of Facebook Groups engagements rather than Forum sessions alone, and we reuse hierarchical, prefix-based semantic IDs (SIDs) learned from cross-platform Facebook Feed data instead of fitting a Forum-specific tokenizer. A 3B-parameter instruction-tuned language model is then supervised-fine-tuned to generate SIDs directly from user context. We systematically ablate the design choices that matter most in practice, including SID construction, the composition and length of user history, and the inclusion of user-profile features. Our results show that cross-platform SIDs transfer to a new recommendation surface, and offer practical guidance for teams deploying GR on real-world social platforms.

cs.IR↗