Search arXivSearch

arXiv · 2608.28668

Reference-Distribution Dependence in LLM-Based Synthetic Persona Data: Diagnosis and Post Hoc Adjustment of Demographic Distributions

Abstract

We diagnose how closely the demographic distributions in LLM-based synthetic persona data match external reference distributions. For the three variables examined, we show that most of the observed error is attributable to the choice of reference rather than to the generator. Using total variation distance (TVD), we compare the sex x age group x province joint distribution of 1,000,000 records from Nemotron-Personas-Korea (NPK) with Korean official statistics. Against resident-registration figures for April 2026, the time of use, the bias bound, defined as the largest possible difference in the share of any subgroup formed from the three variables, is 1.81 percentage points. This is comparable to the margin of error of a survey of roughly 2,900 respondents. This value is not a fixed property of the data. Matching the reference period and series to the generating reference identified here, the 2024 register-based census restricted to Korean nationals, lowers it to 0.56 percentage points. Over the 15 months between the best-matching month (January 2025) and the time of use, the resident-registration population structure itself moves more than twice the distance of NPK's minimum error. Raking and cell post-stratification, the two weighting schemes used in Korean survey practice, remove most of the reference-period dependence at a variance inflation of about 0.2% in both cases. After raking against the generating reference, the residual joint discrepancy lies at, and marginally above, the upper bound of what a perfect generator would produce when realizing 1,000,000 records (97.6th percentile of the Monte Carlo distribution). We recommend treating synthetic persona data as auxiliary material for small-scale survey design rather than as a substitute for survey data, and re-running both diagnosis and adjustment against official statistics current at the time of use.

Explore related subjects

Keep this discovery

BibTeXRIS

Eunjeong Song, Sehee Hong. 2026-08-24. Reference-Distribution Dependence in LLM-Based Synthetic Persona Data: Diagnosis and Post Hoc Adjustment of Demographic Distributions. https://arxiv.org/abs/2608.28668

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related discoveries

Text Data Analysis and Classification Methods - Insights from Customer Letters in Life Insurance

The business of life insurance companies is characterized by long-term contracts. For this reason, data describing customers is of immense value. A portion of the data provided to the customer is rarely or not at all analyzed. This includes customer letters of any kind. This work focuses on classifying customer letters as cancellations and identifying the respective reason, if available. The outlined approach can also be applied to other business transactions and reasons. We discuss data acquisition and preparation, present alternatives, and explain the reasons for the chosen approach. A successful implementation of such a tool can lead to a better understanding of customer cancellation behavior by the insurer, enabling more targeted actions in certain situations.

stat.AP

Measuring the Installed Base: Nordic Health Dataset Catalogues Against HealthDCAT-AP Release 7

The European Health Data Space requires member states to publish machine readable descriptions of the health datasets available for secondary use, and the European Commission publishes HealthDCAT-AP as the metadata profile those descriptions are meant to satisfy. The profile has been designed and validated against curated examples, never against the catalogues already live. We report that measurement for the Nordic region. On 25 August 2026 the 11 Nordic national catalogues harvested by the European data portal held 2,811 dataset descriptions carrying the EU health theme, and none satisfies all eight properties HealthDCAT-AP Release 7 makes mandatory on a dataset. Three of the eight are present on exactly zero records across five countries. Set beside the portal's own quality assessment, which validates DCAT-AP and never mentions the health profile, this is not a health extension skipped on top of sound generic practice: no Nordic catalogue reaches the assessment's top rating band, 5 of the 11 are reported at zero per cent DCAT-AP compliant, and the properties surviving in both layers are the ones a human types into a form, not the ones needing a value bound to a controlled vocabulary. Two further results follow. Finland contributes 2,259 descriptions to the European portal of which 1,146 carry a theme, and not one uses the EU theme authority vocabulary, so a European health filter returns no Finnish dataset at all. Separately, the authority namespace answers HTTP 200 with a well formed empty document for terms it never defined, letting 1,238 datasets across the wider portal carry theme IRIs that resolve to nothing while passing any status code check. We publish the vocabulary, the shapes and the harvesters, record every verdict as a dated observation rather than a property of the dataset, and report the five errors this discipline caught before publication.

cs.DL

Diffusion Distillation for Efficient Weather Ensembles

Diffusion models generate skillful weather ensembles but require costly iterative sampling. We introduce a supervised energy-distance distillation method that compresses a multi-step diffusion teacher into a single-step student by aligning student forecasts with teacher samples and ground-truth observations. Experiments on global forecasting and typhoon-track prediction show that our student outperforms existing distillation methods and preserves skill for extreme events. It matches or surpasses the teacher across key metrics using only one neural function evaluation per autoregressive step.

cs.LG