Source-linked AI summary

Reference-Distribution Dependence in LLM-Based Synthetic Persona Data: Diagnosis and Post Hoc Adjustment of Demographic Distributions

Eunjeong Song, Sehee Hong

arXiv:2608.28668v1cs.CYstat.AP

TL;DR

The paper asks how closely LLM-based synthetic persona demographics match official reference distributions and whether apparent error belongs to the generator or the chosen reference. It compares NPK with Korean statistics using TVD and evaluates reference variation and post hoc weighting. Most observed joint discrepancy is attributable to reference choice, while adjustment leaves residual error near finite-sample variability for the studied variables.

  • Problem

    Synthetic persona research lacks evidence on whether demographic-distribution discrepancies reflect the generator or the external reference, despite demographics forming the base for simulated responses.

  • Method

    The study compares NPK's sex × age group × province distribution with Korean official statistics across reference periods and series, then applies raking and cell post-stratification.

  • Results

    1.81 percentage points against April 2026 resident-registration statistics falls to 0.56 percentage points against the identified 2024 Korean-national census generating reference, with most discrepancy attributable to the reference.

  • Takeaways & Limitations

    Synthetic personas should be treated as auxiliary material for small-scale survey design, with diagnosis and adjustment rerun against statistics current at the time of use.

  • Takeaways & Limitations

    The evidence is a single NPK case limited to sex, age, and province, and the residual cannot separate systematic distortion from one-off realization variability.

Abstract

from arXiv · show

We diagnose how closely the demographic distributions in LLM-based synthetic persona data match external reference distributions. For the three variables examined, we show that most of the observed error is attributable to the choice of reference rather than to the generator. Using total variation distance (TVD), we compare the sex x age group x province joint distribution of 1,000,000 records from Nemotron-Personas-Korea (NPK) with Korean official statistics. Against resident-registration figures for April 2026, the time of use, the bias bound, defined as the largest possible difference in the share of any subgroup formed from the three variables, is 1.81 percentage points. This is comparable to the margin of error of a survey of roughly 2,900 respondents. This value is not a fixed property of the data. Matching the reference period and series to the generating reference identified here, the 2024 register-based census restricted to Korean nationals, lowers it to 0.56 percentage points. Over the 15 months between the best-matching month (January 2025) and the time of use, the resident-registration population structure itself moves more than twice the distance of NPK's minimum error. Raking and cell post-stratification, the two weighting schemes used in Korean survey practice, remove most of the reference-period dependence at a variance inflation of about 0.2% in both cases. After raking against the generating reference, the residual joint discrepancy lies at, and marginally above, the upper bound of what a perfect generator would produce when realizing 1,000,000 records (97.6th percentile of the Monte Carlo distribution). We recommend treating synthetic persona data as auxiliary material for small-scale survey design rather than as a substitute for survey data, and re-running both diagnosis and adjustment against official statistics current at the time of use.

1 Introduction

The paper asks whether synthetic personas reproduce target-population demographics closely enough for survey-related use. It diagnoses this agreement while separating generator discrepancy from dependence on the reference period and series.

  • Sex, age, and region are basic variables for sample design, stratification, weighting, and post hoc adjustment.
  • Joint and marginal distributions of these variables provide a basic quality diagnostic for synthetic persona data.
  • Reference choice can dominate observed discrepancy, because population structure changes between generation and later use.
  • The study varies official-statistics period and series while evaluating NPK's sex, age, and region distributions.
  • RQ2 uses discrepancy measurement to identify the official population that NPK reproduces, whose series and period are undocumented.
  • The paper expresses discrepancy in survey-estimation units, compares it with interpretable baselines, and evaluates two Korean weighting schemes.

2 Background

Synthetic personas provide fictional demographic profiles and language descriptions without real respondents, but their demographic base distribution must be evaluated separately from response plausibility. This paper extends prior synthetic-data audits by varying both reference period and series.

  • Synthetic personas combine structured attributes with natural-language descriptions and do not contain real respondents.
  • They are discussed as possible supplements or replacements for actual surveys, although they are not survey data themselves.
  • Prior criticism has focused on response-generation discrepancies, including variance, regression coefficients, prompt sensitivity, and temporal stability.
  • This study instead examines demographic composition before response generation, treating mismatches in sex, age, and province as coverage error.
  • NPK models attribute relationships with a probabilistic graphical model and generates descriptions with a large language model.
  • The prior audit held each comparison to one period and series, whereas this study varies both axes for a target combination without a documented independence assumption.
  • NPK documentation leaves the generating series and period unspecified and notes timeliness and independence assumptions as relevant constraints.

3 Method

The analysis compares 1,000,000 NPK records with Korean official demographic distributions on a harmonized sex, age, and province system. It uses resident registration as the primary time-of-use reference and census series to investigate the undocumented generating reference.

  • The dataset contains 1,000,000 records, all of which enter the analysis after verification and category mapping.
  • The main variables are sex, age, and province, with analysis restricted to province level because district-level administrative crosswalks are unavailable.
  • Resident-registration statistics as of April 30, 2026 provide the primary external reference for the time-of-use comparison.
  • Resident registration is a de jure administrative count, while the register-based census enumerates usual residents and constitutes a different series.
  • The main age groups are 19–29, 30–39, 40–49, 50–59, 60–69, and 70 and over, with province harmonized to 17 first-level units.
  • All comparisons normalize sources to proportional distributions over a common category system before computing distance.

3.3 Analytic Approach

The analytic approach uses TVD to translate demographic disagreement into survey-relevant bias bounds and compares it across reference distributions and matched baselines. It also evaluates cell post-stratification and raking as post hoc adjustments.

  • Distance and interpretation: The study treats discrepancy as TVD between fixed distributions and interprets it using survey-estimation units rather than significance tests.
  • Distance and interpretation: TVD is half the sum of absolute category-probability differences and equals the maximum event-probability difference between distributions.
  • Distance and interpretation: Multiplying TVD by 100 yields the bias bound in percentage points for any subgroup defined by sex, age group, and province.
  • Distance and interpretation: Interpreting the bias bound as survey-estimate bias assumes identical within-cell response distributions for synthetic and real populations.
  • Baselines: The independence baseline preserves official marginals while removing associations among the three variables, isolating population joint-structure magnitude.
  • Baselines: The multinomial-sampling baseline estimates finite-record TVD variability from 10,000 draws of 1,000,000 records under the official joint distribution.
  • Post hoc adjustment: Cell post-stratification matches every joint cell exactly, whereas raking iteratively matches marginal totals and requires only marginal controls.
  • Reference choice: Reference sensitivity is examined across monthly periods, official series, age groups, and residual discrepancies after adjustment.

4 Results

NPK’s measured demographic discrepancy varies substantially with the official reference period and series, so much of the apparent error reflects reference mismatch rather than generation. Post hoc adjustment removes most period dependence, but residual differences remain tied partly to series definitions and partly to the generator.

  • 4.1 Distributional discrepancy: 0.0159 was the largest marginal TVD, for age group, while age group × province was the largest bivariate discrepancy.The maximum-error marginal category was 70 and over, concentrating discrepancy in combinations involving age.
  • 4.1 Distributional discrepancy: 0.0181 was NPK’s trivariate TVD, equal to 35% of the 0.0520 independence baseline and 3.5 times the multinomial-sampling mean.The observed discrepancy therefore exceeded finite-sample variability alone, although the 35% ratio cannot by itself be interpreted as unreproduced association.
  • 4.1 Distributional discrepancy: 1.81 percentage points was the corresponding bias bound, comparable to the 95% margin of error for approximately 2,900 respondents.The bias bound expresses the maximum possible subgroup-share difference for subgroups formed from sex, age group, and province.
  • 4.2.1 Reference period: 0.0075 to 0.0222 was the TVD range across reference months, with January and February 2025 attaining the minimum.The independence baseline stayed nearly constant at 0.0511 to 0.0520, indicating that reference-month agreement, rather than population association magnitude, drove the change.
  • 4.2.1 Reference period: 15 months separated the best-matching month from April 2026, while resident-registration structure moved 2.1 times NPK’s minimum discrepancy.At April 2026, the discrepancy was 2.4 times the minimum and increased after the minimum by an average of 0.0007 per month.
  • 4.2.2 Reference series: The census Korean-national series matched NPK better than the total-population series, whose minimum shifted to July 2024 and whose curve correlation was only .605.The total-population series includes foreign residents, whereas NPK tracks the Korean-national series.
  • 4.2.2 Reference series: 0.0056 to 0.0142 was the reference-dependent NPK TVD range, reducing the bias bound from 1.81 to 0.56 percentage points against the generating reference.The equivalent sample size rose from approximately 2,900 to approximately 30,000 respondents.
  • 4.2.3 Decomposition by age group: 1.13 percentage points was the 70-and-over age-group error at time of use, falling to 0.21 at January 2025 and no longer being the maximum.At the best-matching month, the largest age-group error was 0.36 percentage points for ages 19–29.

5 Discussion

The discussion shows that reference choice, rather than the generator alone, largely determines measured demographic discrepancy, and that current-reference adjustment can reduce most of it. It also bounds practical use to the diagnosed variables and small-scale survey support.

  • 1.81 percentage points was the time-of-use bias bound against resident-registration statistics, equivalent to a survey margin of error for approximately 2,900 respondents.Matching the identified generating reference reduced the bias bound to 0.56 percentage points.
  • Most discrepancy at the time of use originated in the reference distribution rather than the generative model.The dataset reproduced the generating reference closely, although its series and period were undocumented.
  • Reference-period drift made external agreement deteriorate by about 0.07 percentage points per month, increasing the bias bound from 0.75 to 1.81 percentage points within 15 months.Diagnosis and adjustment should therefore be rerun against official statistics current at the time of use.
  • Cell post-stratification removes the joint discrepancy completely at 0.21% variance inflation, while raking remains an alternative when joint totals are unavailable or cells are sparse.Raking leaves a residual of about 0.007, and weighting implications should be reevaluated for subsamples.
  • The diagnostic supports sex, 10-year age group, and province use, but finer age resolution leaves discrepancies and empty cells that can prevent cell post-stratification.Fine-grained subgroup scenarios require rerunning the diagnostic on the reduced grid.
  • The case supports auxiliary use in small-scale survey design, not replacing large surveys, estimating population parameters, public opinion, or trends.Bias does not decrease with survey size, and the diagnostic covers only demographic variables.

6 Conclusion

The conclusion finds that demographic discrepancy in NPK depends strongly on the official reference period and series. It recommends current-statistics diagnosis and adjustment while treating synthetic personas as auxiliary material for small-scale survey design.

  • Reference choice changed the observed discrepancy by nearly threefold across months and 2.5-fold across series.
  • Synthetic persona data should be treated as auxiliary material for small-scale survey design rather than as a substitute for survey data.
  • Diagnosis and adjustment should be rerun against official statistics current at the time of use.

Funding

The study reports that it received no funding.

  • No funding was received for this study.

AI Use Disclosure

The authors used DeepL during manuscript preparation for translation and English-language editing, then reviewed and edited the output.

  • DeepL was used to translate the initial draft and improve English readability.The authors reviewed and edited the output and take responsibility for the final content.

Appendix A: Reference Months Used in the Reference-Period Analysis

Appendix A lists the resident-registration reference months used in the reference-period analysis, with monthly refinement used to locate the NPK distance minimum.

  • The reference-month table covers the resident-registration series used in the reference-period analysis.
  • The main grid uses six-month spacing from April 2023 to April 2024 and three-month spacing thereafter.
  • Additional monthly reference months were used solely to locate the minimum of the NPK curve.

Appendix B: Province Category Crosswalk

Appendix B provides a table mapping synthetic-data province names to official categories.

  • Table B1 maps synthetic-data province names to official categories.

NPK abbreviated name

The appendix identifies abbreviated and official province names alongside English labels used in the analysis, with source-file naming variants normalized for matching.

  • The province table pairs abbreviated names with official Korean names and English labels used here.
  • The listed categories include Gangwon, North and South Chungcheong, North and South Jeolla, North and South Gyeongsang, and Jeju.
  • Province-name variants were normalized so spacing and longer source-file forms map to the same categories.
Loading 2608.28668v1…