arXiv · 2609.33666
One Model Is Not a Crowd: Multi-LLM and Aspect-Conditioned Diverse Comment Generation
Abstract
Human communication on the internet is shaped by diverse perspectives, most visibly expressed in online comment spaces. As large language model (LLM)based AI agents begin to inhabit these spaces, a key question arises: whether synthetic comment threads can capture the diversity inherent in human discourse. This concern is increasingly important, as the growing presence of homogenized AI-generated content risks reducing diversity over time, potentially leading to model collapse and degrading the richness of digital communication. Inspired by the plurality of human crowds and the aspect-driven nature of discourse, we hypothesize that comment diversity is better approximated by combining multiple LLMs with aspect-conditioned generation. We formalize and evaluate this approach using models from different providers and introduce a framework that characterizes diversity across semantic, linguistic, and socio-pragmatic features along three axes: dispersion, coverage, and alignment. Using this framework, we conduct a large-scale study on over 2 million YouTube comments across multiple domains. Our results reveal that multi-LLM and aspect-conditioned generation better align with human comment distributions and such data remains viable under pretraining style curation and is effective for downstream tasks. Yet, human diversity remains unmatched. Overall, our findings provide a practical foundation for generating more diverse and socially grounded discourse in AI-mediated environments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nafis Irtiza Tripto, Delvin Ce Zhang, Mahjabin Nahar, Dongwon Lee. 2026-09-27. One Model Is Not a Crowd: Multi-LLM and Aspect-Conditioned Diverse Comment Generation. https://arxiv.org/abs/2609.33666
Cite the original work for its findings. Save a collection to share your selection of sources.