Source-linked AI summary

Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey

Howard Kim, Keun Tae Cho

arXiv:2608.28615v1cs.CYcs.CL

TL;DR

Synthetic survey respondents have not been systematically validated outside English-speaking contexts, especially for Korean digital and AI service use. This study compares Korean LLM-based personas with weighted Korea Media Panel Survey estimates and finds limited, model-specific validity: calibration helps within-wave, but real-data estimators remain more accurate except when real data are nearly absent.

  • Problem

    Whether synthetic responses reliably reproduce real population distributions remains insufficiently validated in non-English contexts such as Korea.

  • Method

    The study evaluates sex-and-age-stratified Korean synthetic persona panels against weighted Korea Media Panel Survey microdata across digital and AI service-use measures.

  • Results

    Synthetic panels show 15–19 pp overall and segment MAE with systematic model-specific bias; calibration roughly halves sex-by-age error, but direct real-data estimation reaches 3.6 pp versus 8.6/6.7 pp.

  • Takeaways & Limitations

    Synthetic persona panels are not substitutes for real measurement; their defensible value is primarily diagnostic and operational use is limited to settings with nearly absent real data or unobserved segments.

  • Takeaways & Limitations

    Generalization beyond Korea, digital and AI service use, the two evaluated LLMs, and the 2024–2025 period remains an open question, while calibration does not transfer across waves.

Abstract

from arXiv · show

Synthetic personas based on large language models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet systematic validation outside English-speaking contexts remains scarce. This secondary-data study evaluates how well a Korean synthetic persona panel (NVIDIA Nemotron-Personas-Korea), conditioned into Gemini 3.5 Flash (primary) and EXAONE (comparison), reproduces digital and AI service-use distributions from the KISDI Korea Media Panel Survey. Sex-and-age-stratified panels of about 8,000 personas per model answered the survey's own items - eight service-use indicators and eight innovativeness and acceptance constructs - and were compared against weighted survey estimates. The overall mean absolute error (MAE; RQ1) was 15-19 percentage points (pp), with binary item-mean correlations of 0.69-0.90. Segment error (RQ2) across five demographic axes was 15-19 pp, with between-group gaps up to 52.4/36.2 pp (Gemini/EXAONE). Errors followed model-specific signatures: an age stereotype with low anchoring (Gemini) versus an acquiescence-consistent level bias (EXAONE). Reference-year analysis was consistent with temporal misalignment driving most generative-AI overestimation, whereas short-form underestimation was framing-sensitive. Holdout calibration on 30% of the real data (RQ3) roughly halved sex-by-age cell MAE (18.9->8.6, 15.9->6.7 pp) - yet direct estimation from the same real subsample was far more accurate (3.6 pp), and the correction did not transfer across time. The calibrated panel retained an advantage only under extremely scarce real data (about 100 responses) and, for one model, for unobserved segments. Persona-narrative conditioning beat demographic-only conditioning, but neither surpassed simple real-data baselines. Synthetic panels are thus not survey substitutes; their value is diagnostic, with operational use confined to settings lacking real data.

I. INTRODUCTION

The study examines whether LLM-conditioned Korean synthetic personas can reproduce real digital and AI service-use distributions, addressing limited validation outside English-speaking contexts. It evaluates overall agreement, demographic segment error, and calibration against the Korea Media Panel Survey.

  • Population surveys are costly and response rates are declining as policymakers need timely measures of rapidly changing digital behaviors.
  • Systematic validation of synthetic responses against real population distributions remains limited, particularly in non-English contexts such as Korea.
  • RQ1 measures overall distributional agreement, RQ2 decomposes error across demographic segments, and RQ3 tests small-data calibration and transfer across waves.
  • Related work: Prior work found reasonable population-level opinion matching from demographic backstories, but later research identified skew, variance compression, temporal and group biases, and response-style problems.
  • Related work: The study evaluates segment fidelity across five demographic axes and reports a corresponding between-group bias index rather than relying only on aggregate means.
  • Related work: Small-sample calibration is positioned as a bias-mitigation strategy, while narrative conditioning is examined as an alternative to demographic-only conditioning.

III. METHOD

The study samples synthetic personas, generates questionnaire responses through two LLMs, and validates post-stratified estimates against weighted KISDI survey estimates. Diagnostics then guide signature-matched calibration evaluated on held-out and later-wave data.

  • Personas are sampled by strata and answer the survey through two LLMs before validation, diagnostics, and signature-matched calibration.
  • Data: The reference is the nationwide Korea Media Panel Survey, while Nemotron-Personas-Korea supplies about one million synthetic persona records.
  • Data: The study measures eight binary digital-service indicators and eight 5-point innovativeness and technology-acceptance constructs.

B. SAMPLE DESIGN

The panel is stratified by sex and age, with controlled per-cell sampling and model comparison designed to assess Korean response fidelity and reproducibility. Persona narratives and generation procedures provide the conditioning inputs.

  • Age is divided into seven decade groups and crossed with sex to form 14 sampling cells.
  • Each cell receives n = clip(nactual × 2, 200, 600) personas, sampled uniformly with replacement to balance precision and cost.
  • The real reference samples contain 8,693 individuals in 2024 and 8,411 in 2025, while valid synthetic samples contain roughly 8,000 responses per model.
  • Gemini 3.5 Flash is the primary closed-weight model and K-EXAONE-236B-A23B is the Korean-developed open-weight comparison model.
  • Each persona combines a first-person summary with cultural background, hobbies, and arts/media orientations that provide relevant media-consumption cues.
  • Responses failing format validation are regenerated up to three times, conditional items follow survey skip logic, and exact hosted-model identifiers and timestamps are logged.

E. VALIDATION METRICS

Validation compares post-stratified synthetic estimates with weighted real estimates using distributional, correlation, divergence, and segment-error metrics. Bootstrap intervals and practical percentage-point magnitudes frame uncertainty and interpretation.

  • RQ1 evaluates MAE, cosine similarity, KL divergence, JS divergence, and item-mean Pearson correlation between synthetic and weighted real estimates.
  • The post-stratified estimate weights each sex-by-age cell’s synthetic mean by that cell’s weighted share of the real sample.
  • Binary indicators are compared on their 0–1 scale, while constructs are evaluated separately on their native 1–5 scale to avoid pooled-correlation distortion.
  • DPDe is the maximum minus minimum signed group error across 14 sex-by-age cells, averaged across indicators.
  • A design-based paired bootstrap with B = 600 resamples households for real data and rows for synthetic personas to produce 95% percentile confidence intervals.
  • Because weighting makes small differences statistically detectable, the analysis emphasizes practical percentage-point magnitude and headline confidence intervals without formal multiplicity correction.

F. CALIBRATION DESIGN (RQ3)

The calibration design learns synthetic-panel bias from 30% of real 2024 data, then evaluates corrections against held-out data and real-data-only baselines.

  • The 2024 real data were split within sex-by-age strata into 30% calibration (n = 2,602) and 70% test (n = 6,091) sets.
  • Calibration estimated synthetic per-cell bias using either a global additive shift or a demographic age regression.The age regression matched the correction form to the observed age-structured bias signature.
  • Real-data-only benchmarks included direct weighted cell-rate estimation, an age–sex regression, and the calibration grand mean.Empty direct-estimation cells fell back to the calibration grand mean.
  • Additional analyses compared grand-mean and prior-wave baselines, demographic-only conditioning, response variance, independent EXAONE passes, and short-form wording variants.Prior-wave comparisons covered six indicators because generative-AI and short-form items were absent in 2023.

IV. RESULTS

Synthetic panels showed substantial overall and demographic-segment disagreement with weighted 2024 survey estimates, despite some agreement in the ordering of behaviors.

  • A. RQ1: OVERALL AGREEMENT: 17.1 pp Gemini MAE and 15.0 pp EXAONE MAE showed substantial overall disagreement with 2024 reference estimates.EXAONE was significantly closer because the design-based 95% CIs did not overlap; sampling error was below 1 pp per indicator.
  • A. RQ1: OVERALL AGREEMENT: 0.795/0.903 binary item-mean correlations and 0.944/0.968 cosine similarities indicated broadly correct behavioral ordering but did not eliminate magnitude errors.KL divergence was 0.116/0.072, favoring EXAONE; correlations covered only eight items and were treated as descriptive.
  • B. RQ2: SEGMENT ERROR: 16.7–18.8 pp Gemini and 14.9–16.9 pp EXAONE five-axis segment MAE showed that demographic disagreement remained substantial.Gemini performed worst on the age axis.
  • B. RQ2: SEGMENT ERROR: 52.4 pp Gemini and 36.2 pp EXAONE DPDe showed large between-group error spreads, with Gemini’s spread potentially reversing segment-level conclusions.Combined sex-by-age cell MAE was 18.9 pp for Gemini and 15.9 pp for EXAONE.
  • B. RQ2: SEGMENT ERROR: Excluding both teen cells preserved the main conclusions, including segment-error patterns, calibration and baseline orderings.Gemini’s overall RQ1 MAE slightly worsened from 17.1 to 17.8 pp.

C. STRUCTURAL BIAS SIGNATURES

The models exhibited distinct structural biases: Gemini amplified an age-linked digital stereotype, while short-form errors were strongly sensitive to item framing and generative-AI errors aligned with reference-year mismatch.

  • C. STRUCTURAL BIAS SIGNATURES: Gemini’s signed-error age slopes were −12.4, −8.6, −7.9, and −7.1 pp per age step for OTT, short-form, SNS, and subscription indicators.These slopes were consistent with an amplified “digital = young” stereotype.
  • C. STRUCTURAL BIAS SIGNATURES: Gemini’s 5-point mean was 2.45 versus the real weighted mean of 2.83, while EXAONE showed a 3.05 mean and positive level bias.EXAONE’s binary yes-rate was 60.2% versus Gemini’s 54.9%.
  • D. ITEM FRAMING: Changing only the short-form wording to remove “OTT” and name venues recovered short-form use to near-real levels.The paired wording-controlled regeneration used approximately 1,100 personas per model and produced changes whose 95% CIs excluded zero.
  • E. TEMPORAL MISMATCH: A 2024-to-2025 reference-year swap shrank the generative-AI gap from +26.2 to +7.3 pp for Gemini and +23.4 to +4.9 pp for EXAONE.The same swap widened short-form gaps, supporting temporal mismatch for AI overestimation but framing-linked short-form error.
  • E. TEMPORAL MISMATCH: Matched-wave regeneration kept generative-AI gaps small but short-form collapsed to −72.2/−57.7 pp, leaving overall MAE unimproved.Excluding short-form would reduce 2025 MAE to 11.8/10.9 pp for Gemini/EXAONE.

F. ROBUSTNESS

Calibration substantially reduces within-wave sex-by-age error, but its gains are modest across time and do not generally outperform direct real-data estimation. Correction effectiveness depends on the model’s bias signature and the availability of real data.

  • 0.5 pp was Gemini’s temperature sensitivity, versus 4.0 pp for EXAONE, indicating greater response sensitivity for EXAONE.
  • 18.9/15.9 pp uncorrected sex-by-age MAE fell to 8.6/6.7 pp after age-regression calibration for Gemini/EXAONE.The corresponding improvements were 10.3 pp and 9.2 pp across repeated within-wave splits.
  • Gemini benefited mainly from the age term, whereas EXAONE gained most from the global correction because their bias signatures differed.For Gemini, global correction changed MAE from 18.9 to 14.6 pp; for EXAONE, it changed 15.9 to 8.8 pp before the age term.
  • 3.6 pp was the direct-estimation error from the 30% real calibration subsample, outperforming calibrated synthetic panels and real-only age–sex regression.The calibrated panel retained a clear advantage only at 1% of data, approximately 87 responses, for EXAONE.
  • 14.5/14.0 pp remained the temporally calibrated Gemini/EXAONE errors, above the 2025 grand-mean and prior-wave references.The temporal analysis indicates that within-wave error reduction did not establish forward-predictive validity across waves.

H. COMPARISON TO NAIVE BASELINES

Synthetic panels perform worse than simple baselines before calibration and remain inferior to a prior-wave baseline afterward. Their relative value is limited to newer indicators or settings with very sparse or unobserved real segments.

  • 18.9/15.9 pp uncorrected synthetic cell MAE exceeded the 11.6 pp grand-mean baseline for Gemini/EXAONE.The persona-imposed segment structure was therefore less accurate than assuming no segment variation.
  • 8.6/6.7 pp calibrated synthetic MAE surpassed the 11.6 pp grand mean for Gemini/EXAONE, but this improvement required calibration.
  • 3.7 pp was the prior-wave error on the six-indicator common set, compared with 7.5/7.0 pp for calibrated Gemini/EXAONE.The prior-wave baseline does not provide estimates for generative-AI or short-form indicators.

I. PERSONA CONDITIONING ABLATION

Narrative persona conditioning improved both models over demographic-only conditioning, but synthetic validity remained imperfect and model-specific. Errors were partly diagnosable and correctable within a wave, while real-data baselines remained stronger except under severe data scarcity.

  • Persona conditioning ablation: 18.9→22.3 pp and 15.9→22.8 pp: demographic-only conditioning increased sex-by-age cell MAE for Gemini and EXAONE, respectively.Overall RQ1 MAE also rose from 17.1 to 18.9 pp for Gemini and from 15.0 to 22.5 pp for EXAONE.
  • Validity and bias: 15–19 pp: synthetic panels reproduced target distributions only imperfectly, with larger subgroup error that decomposed into identifiable causes.The structured errors were diagnosable and partially correctable within a survey wave.
  • Temporal alignment: Reference-year alignment removed most generative-AI overestimation, but overall MAE did not improve because temporal misalignment explained only specific indicators.Short-form use remained underestimated, with the gap widening as real adoption rose.
  • Model-specific bias signatures: Gemini showed low anchoring and an age stereotype, whereas EXAONE showed an acquiescence-consistent upward level shift.The comparison is descriptive because the systems differ in scale, training data, and serving stack.
  • Framing sensitivity: The OTT-service framing caused synthetic short-form endorsement to collapse despite accurate YouTube estimates.This makes wording standardization and reporting important for synthetic validity.
  • Calibration and operational value: 3.6 vs 8.6/6.7 pp: direct estimation from the same real calibration subsample was more accurate than the calibrated panel at the 30% fraction.The calibrated panel retained an edge only around n≈100 or for entirely unobserved segments, with the latter only for EXAONE.
  • Practical implication: Beyond a few hundred real respondents, direct estimation is more accurate, so synthetic panels should be reserved for nearly absent data or uncovered segments.Operational use should be benchmarked against real-data-only estimators before deployment.

VI. CONCLUSION AND LIMITATIONS

Across the three research questions, the Korean synthetic panels showed limited, systematic validity and only narrow calibration gains. Their defensible role is diagnostic rather than substitutive, especially because calibration did not transfer across waves.

  • RQ1: Overall agreement: 15–19 pp: overall agreement was limited, with systematic model-specific bias rather than random noise.RQ1 assessed agreement between synthetic and real response distributions.
  • RQ2: Segment error: 52.4 pp and 36.2 pp: Gemini and EXAONE showed sharply different maximum between-group error gaps, respectively.Error concentrated especially among older adults under Gemini.
  • RQ3: Calibration: Calibration roughly halved sex-by-age cell error, but direct estimation was more accurate beyond a few hundred real responses and correction failed across waves.Out-of-time calibration trailed both grand-mean and prior-wave references.
  • Implications and scope: Synthetic persona panels are not substitutes for real measurement; defensible uses are diagnostic and operationally limited to nearly absent data or entirely unobserved segments.The stated scope is Korea, digital and AI service use, two LLMs, and the 2024–2025 window.

APPENDIX: SENSITIVITY AND SIGNED ERROR BY AGE BAND

Sensitivity analyses preserve the headline conclusions after excluding the least comparable teen cells. Signed age-band errors reveal Gemini’s steep age gradients and EXAONE’s flatter level shift, within a secondary analysis using de-identified data.

  • Sensitivity analysis: 12 cells: the teen-cell-excluded sensitivity analysis preserved all orderings and conclusions while shrinking DPDe without changing headline error levels.The excluded cells were the teen cells considered least population-comparable.
  • Appendix tables: Tables 8 and 9 report 2024 signed errors by age band for Gemini and EXAONE, respectively.The errors are defined as synthetic minus actual, in percentage points, with sexes pooled.
  • Gemini signed error: +71 pp to +2 pp: Gemini’s generative-AI overestimation declined from teens to people in their 70s-and-over.Its OTT, short-form, and SNS errors deepened through middle and older bands before easing somewhat in the oldest band.
  • EXAONE signed error: EXAONE’s errors were flatter across age but shifted in level.The age gradients were fitted on 14 sex-by-age cells, so pooled-band slopes can differ slightly.
  • Study status: The study used de-identified KISDI microdata and synthetic persona responses, without recruiting new human subjects.No personally identifiable information was used or released.
  • Reproducibility: The reproducibility package archives code, prompts, logs, seeds, raw responses, weighting and calibration code, and preregistration correspondence.The package is publicly archived at the repository URL specified in the paper.
Loading 2608.28615v1…