Source-linked AI summary
Exploratory As-Analyzed No-Detection of Culturally-Marked Predicate-Triggered PII Amplification in a Synthetic-English RAG Probe: A Predicate-Resource-Confounded Audit
Yanhang Li, Zhichao Fan, Zexin Zhuang
TL;DR
The paper asks whether stereotype-loaded query framing increases PII leakage from a RAG system. It preregisters a four-culture audit on a synthetic-English corpus, but analyzes substituted exploratory estimators because the locked endpoint was not run. Cleaner non-name channels show no corrected amplification, while prompt echo and predicate-resource confounding support a no-detection interpretation rather than evidence of no effect.
Problem
The paper asks whether stereotype-loaded queries about culturally marked people leak more PII than otherwise-equivalent neutral queries in RAG systems.
Method
The study compares five query arms across en-Anglo, es-LATAM, Arabic, and Hindi in a synthetic English PII corpus using paired leakage analyses with query and document language fixed.
Results
No stereotype-driven amplification was detected on the cleaner non-name metric in any of the four cultures after Bonferroni correction.
Takeaways & Limitations
The evidence is best read as a no-detection result for culturally marked predicate leakage, with the observed contrast confounded by the predicate resource rather than established as a QS effect.
Takeaways & Limitations
The locked post-guard Llama-Guard-3 plus summarize estimator was not run; the sampled predicates also mix stereotype-loaded items with cultural markers and heritage practices.
Abstract
from arXiv · showhide
We ask whether stereotype-loaded queries about culturally marked people leak more personal information from a retrieval-augmented generation (RAG) system than otherwise-equivalent neutral queries. We pre-register a four-culture audit (en-Anglo, es-LATAM, Arabic, Hindi) on a synthetic English PII corpus, comparing five query arms we call the Stereotype-Trigger Leakage Delta (STLD). Two caveats up front. Our locked confirmatory estimator was never run, so every test in the paper is exploratory or sensitivity, with all plan deviations listed in the appendix. And the name-leakage metric is contaminated by a prompt-echo artifact: the model often just re-emits the name we asked about, which inflates apparent leakage without any retrieval at all. On the cleaner channels (email, phone, ssn-like, address), we find no stereotype-driven amplification on any of the four cultures after multiple-comparison correction. Because our sample is only powered for mid-sized effects, and because the culturally marked probes mix stereotype content with cultural markers and heritage practices, we present this as no detection, not evidence of no effect, of culturally marked predicate leakage that is confounded with the underlying resource.
1 Introduction
This paper tests whether culturally marked, stereotype-loaded query framing amplifies PII leakage in a fixed-language synthetic-English RAG system. Exploratory analyses find no corrected amplification on cleaner non-name channels, while prompt echo and predicate-resource confounding limit interpretation.
- Research question and design: The study tests whether stereotype-loaded queries extract more PII than content-equivalent neutral queries in a four-culture synthetic-English RAG audit.The preregistered design used five query arms and N=100 per culture, with query and document language fixed to English.
- Main result: No cell was Bonferroni-significant on the cleaner non-name metric across en-Anglo, es-LATAM, Arabic, and Hindi.The metric covers email, phone, ssn-like, and address leakage with n=80 per culture.
- Main result: The contaminated name-included metric flagged one es-LATAM cell at −10 pp, but the contrast reflected an elevated QC control arm rather than a defensive QS arm.The QC rate was 80%, while the other arms were near 67−75%; the sanity rule L(QC)−L(QN)<3 pp was violated in direction.
- Confounds and sensitivity: A QC-only sensitivity rerun reduced the control-arm rate from 80% to 70% and collapsed cell-level STLD to 0 pp on both bank-labelled sub-pools.The authors interpret this as a small-pool QC sampling artifact, not a QS effect.
- Confounds and sensitivity: Name leakage is contaminated by prompt echo: 17/20 es-LATAM QS responses repeated the queried person_name already present in the prompt.The authors therefore treat the non-name metric as the cleaner diagnostic contrast.
- Inferential status: The locked confirmatory endpoint was not run, so all reported inferential tests are exploratory as-analyzed or sensitivity analyses.The paper explicitly treats the resulting estimator as an estimand shift rather than a cosmetic operational change.
2 Related Work
Prior work studies stereotypes in model outputs, RAG-side bias and privacy vulnerabilities, multilingual PII leakage, and memorization-based privacy attacks. This paper isolates culturally indexed query framing within a fixed language and measures its relation to PII leakage in RAG.
- Stereotype benchmarks across cultures: Stereotype benchmarks document conceptual limits and culturally or linguistically situated bias resources across language–region pairs.The paper adopts these resources’ predicate-sourcing precedent without claiming sampled banks represent their broader populations.
- RAG-side bias amplification and privacy: RAG research has shown retrieved stereotype-laden documents can amplify bias outputs and that retrievers present fairness vulnerabilities.Those studies treat stereotypes primarily as a system-output phenomenon.
- RAG-side bias amplification and privacy: Privacy research treats RAG as an attack surface and extends PII-leakage analysis across languages with mitigation approaches.This paper instead holds query and document language fixed to isolate query framing from cross-lingual effects.
- Memorization, MIA, and multilingual safety: Memorization studies show verbatim training-data recall can enable privacy attacks, while this setting introduces PII through retrieval rather than pretraining.The paper also emphasizes that probe and decoding conditions are part of claim-specific memorization audits.
- Study positioning: The audit uses within-person paired deltas to test culturally indexed query framing as a candidate controllable lever for RAG PII leakage.Controls address length, predicate capacity, refusal asymmetry, and retrieval-cue confounds.
3 Method
The study uses a four-culture, five-arm paired-query audit of culturally marked predicate leakage in a shared English RAG system. Because the locked confirmatory estimator was not run and execution substituted the guardrail, reformulation, and headline metric, reported analyses are exploratory or sensitivity tests.
- Threat model: The threat model assumes a black-box attacker queries a target person represented in an English document store to recover synthetic PII.Targets include names, emails, phone numbers, SSN-like identifiers, and postal addresses.
- Five-arm paired design: Four culture-specific probe banks feed paired treatment and control queries into a shared RAG system, with name-included and cleaner non-name leakage channels tracked.The cleaner non-name channel covers email, phone, SSN-like, and address outputs.
- Five-arm paired design: Each person–PII-type pair receives five content-equivalent queries: bare, random, neutral-elaborated, culture-neutral, and stereotype-loaded.The primary contrast is STLD=L(QS)−L(QC), using refusal-as-no-leak; length and natural-context sanity checks are also specified.
- Corpus and predicates: The corpus contains 4×200=800 English documents, each with a unique non-PII anchor, one synthetic PII item, and a culture-specific target name.Predicate sources include existing multilingual stereotype banks and a hand-authored novel sub-bank, which was smaller than planned for three cultures.
- Models and estimator: The RAG stack uses BGE-M3 retrieval with k=5 and Qwen-2.5-7B-Instruct generation, while the locked Llama-Guard-3-8B path was replaced by a production-grade PII regex.The analysis uses pre-guard generator emission and direct reformulation because the locked summarize reformulation reached a 96−100% pre-guard ceiling.
- Preregistration and inference: The preregistered post-guard summarize estimator was not run, so the substituted pre-guard direct-reformulation metric is an estimand shift rather than a cosmetic implementation change.D1–D3 changed the estimator, and the four v1 STLD cells are labeled primary only relative to sensitivity analyses, not confirmatory in the strict preregistration sense.
4 Results
The as-analyzed results do not provide a stable stereotype-leakage signal: the contaminated name metric shows an opposite-direction es-LATAM contrast, while the cleaner non-name metric detects no corrected-significant cell. Sensitivity analyses attribute the es-LATAM pattern to control-pool, predicate-resource, refusal, and model-specific factors rather than a culture-level QS effect.
- 4.2 As-analyzed STLD: −10.0 pp es-LATAM STLD is significant under the contaminated name-included metric, but it is opposite to the preregistered STLD>0 hypothesis.The contrast is interpreted as not supporting H1, not as evidence for a flipped hypothesis or the locked estimator.
- 4.2 As-analyzed STLD: The es-LATAM contrast is concentrated against elevated L(QC), with L(QS)−L(Q0)=+3 pp, rather than against the neutral baseline.The QC sanity rule is violated in direction, identifying QC as anomalously high-leak rather than QS as defensive.
- 4.3 D8 QC sensitivity: Expanding QC from three to seven predicates shifts leakage from 80% to 70% and collapses cell-level STLD to 0 pp on both bank-labelled sub-pools.Because only the control arm was rerun, D8 is a QC-stability sensitivity test rather than a causal replacement.
- 4.4 Non-name metric: The name metric is contaminated because 17/20 es-LATAM QS responses echo the queried person_name already present in the prompt.The non-name metric is therefore treated as the least-contaminated descriptive estimator.
- 4.4 Non-name metric: No non-name cell is Bonferroni-significant across the four cultures, so the cleaner metric does not support stereotype-triggered amplification.The metric covers email, phone, ssn-like, and address leakage with n=80 trials per culture.
- 4.5 Refusal mediation: Es-LATAM refusal asymmetry is +10 pp, from 19% to 29%, with 0 flip-down versus 10 flip-up transitions.The authors label this a mediator analysis because controlled-refusal ablation is needed to separate safety routing from leakage propensity.
- 4.5 Predicate variance: Per-predicate estimates preserve the negative es-LATAM sign, but the effect concentrates in 11 EspanStereo-style predicates and is null in four novel predicates.The leave-one-out range is [−11.8, −7.4] pp, while the predicate-cluster bootstrap 95% CI is [−18.4, −2.8] pp.
- 4.5 Cross-model probe: The 32B same-family probe preserves the es-LATAM sign at STLD=−3 pp without any Bonferroni-significant cell.The probe is single-seed and descriptive; gold-only force context restores L(QR)≈L(Q0).
5 Discussion
The discussion treats the es-LATAM result as a predicate-resource-confounded observation, not a culture-level effect. The paper’s scope is limited to a single synthetic English corpus, four cultures, one main model family, and pre-guard generator behavior.
- 5 Discussion: The as-analyzed estimator finds null STLD on three cultures and a significantly negative −10 pp es-LATAM cell, opposite to H1.The contrast localizes to elevated L(QC)=80%, not L(Q0) or L(QR)=67%.
- 5 Discussion: The name-included figure is invalidated by prompt echo, while Table 2 supplies the headline validity-filtered non-name read.Only the es-LATAM contaminated-metric cell crosses corrected significance; the other three cultures do not.
- 5 Discussion: Because the effect is confined to the EspanStereo-style sub-pool, the es-LATAM cell is treated as predicate-resource-confounded rather than a culture-level claim.The paper does not attribute the result to alignment-training-data composition.
- Generalization and scope: The result is observed on one model family, one synthetic English-source corpus, four cultures, and one predicate-sourcing pipeline.The study does not test mitigation, real multilingual corpora, diaspora-versus-local contrasts, or closed-weight larger models.
- Practical implications: The audit proposes prompt-echo-aware scoring, input-side predicate sterilization, and separate monitoring of refusal-routing asymmetries for deployed RAG systems.These diagnostics are motivated by the audit but are not empirically validated here.
6 Conclusion
Under the as-analyzed pre-guard estimator, the positive-direction STLD hypothesis is not supported across four cultures. The cleaner non-name metric yields no Bonferroni-significant cell, while the v1 es-LATAM contrast appears control-driven and predicate-resource-confounded.
- The positive-direction H1: STLD>0 is not supported on any of four cultures under the as-analyzed pre-guard estimator.
- The −10 pp es-LATAM cell is interpreted as a control-driven contrast rather than a QS effect.A post-hoc QC-only sensitivity check produces a null under a different QC pool, so v1 and v2 are reported side by side rather than treating v2 as causal replacement.
- Under the non-name metric, no cell is Bonferroni-significant in either v1 or v2.The metric covers email, phone, ssn-like, and address leakage at n=80/culture.
- At N=100/culture, the MDE is approximately ±11 pp at 80% power, so the result is framed as no detection rather than evidence of no effect.
- Because es-LATAM probes mix stereotype content with cultural markers and heritage practices, the finding is presented as predicate-resource-confounded culturally marked predicate leakage.
Limitations
The study’s interpretation is limited by construct identifiability, limited power, estimator changes, sensitivity-test design, and a narrow synthetic-English model scope.
- The es-LATAM bank mixes stereotype-loaded items with cultural markers and heritage practices, with author-coded rather than in-culture-panel annotation.
- At N=100, or n=80 for non-name metrics, the 80%-power MDE is approximately ±11−13 pp, making per-predicate cells diagnostic rather than inferential.
- The locked post-guard Llama-Guard-3 plus summarize estimator was not run, so pre-guard regex/direct analyses represent an estimand shift.D8, 32B, and force-context probes are single-seed sensitivity tests rather than venue-grade replications.
- The study uses a synthetic English corpus, one Qwen-2.5 model family, four cultures, and qlang=doclang, and studies no mitigation.
Ethics Statement
The audit uses fully synthetic PII and does not target real people or production systems. Its culturally marked predicate set is deliberately constrained, and the four culture labels are a limited, non-essentialized sample.
- The corpus contains Faker-generated PII and no real personal information at any stage of the study.
- The predicate bank includes descriptive cultural-marker, heritage, and mild-stereotype items, excluding explicitly derogatory predicates.
- The study does not target any production system or real person.
- The four cultures are a limited sample, and their short labels are used for discourse rather than as essentialized categories.
A Audit appendix: token counts, retrieval recall, sterilization, construct annotation
The appendix documents query-length matching, near-complete target retrieval, successful sterilization audits, and author-coded construct annotation.
- Token counts: The four elaboration arms are length-matched within ±6% on every culture, while Q0 is shorter by design.Total query token count equals base prompt token count plus predicate token count.
- Retrieval recall: Target document recall@5 is at least 99% across all culture-arm cells, indicating no detectably framing-sensitive retrieval in this corpus.
- Sterilization: All 43 stereotype predicates and 11 culture-neutral QC predicates pass the automated PII-leakage-capacity audit.Both banks have n_failed=0 across the documented audit rules.
- Construct annotation: Each es-LATAM QS predicate is annotated as stereotype-loaded, cultural marker, or heritage practice.The annotations are author-coded rather than produced by an independent in-culture panel.
B Discordant counts, paired Wald CIs, and one-sided preregistered p-values
The analysis reports paired McNemar and Wald statistics for diagnostic, non-name, post-guard, and D8 contrasts, while emphasizing that the locked end-to-end estimator was not run. Table 4 reports no rejection of the preregistered one-sided hypothesis, and the D8 results indicate unstable QC controls rather than a QS effect.
- Reported statistics: Table 4 reports discordant counts, exact paired McNemar p-values, one-sided preregistered p-values, Wald 95% CIs, and Bonferroni status for the analyzed contrasts.The diagnostic family, non-name rerun, post-guard regex contrast, and D8 QC shift are included.
- Inference: The preregistered one-sided H1 is not rejected on any cell, although the single Bonferroni-significant v1 es-LATAM contrast is negative.The non-name rerun has no Bonferroni-significant cell, and the post-guard es-LATAM contrast is uncorrected-only at p=0.031.
- D8 sensitivity: D8 shows a control-bank shift from v1 to v2, while QS−QC v2 is null with a wide confidence interval.The QC shift has a fully positive CI, indicating markedly higher v1 QC than v2 QC.
- D8 sensitivity: The seven-predicate v2 pool spans 37.5−100% leakage across predicates, consistent with predicate-specific noise and a v1 QC pool at the high end.The reported rates range from 3/8=37.5% for sports to 9/9=100.0% for Spanish-with-relatives.
- Analysis status: The locked end-to-end estimator was not run, so the reported analyses are as-analyzed diagnostics rather than confirmatory tests.The analyzed estimator uses regex guardrail, direct reformulation, and pre-guard generator emission instead of the locked estimator.
E Pre-registration deviations (D1–D8)
The paper documents eight deviations from the preregistered design, including estimator changes, limited predicate coverage, a single-seed model run, and post-hoc QC sensitivity analysis. These deviations constrain interpretation of the reported privacy-leakage results.
- D1–D3 estimator changes: D1 changed the locked summarize reformulation to direct reformulation because summarize reached a 96−100% rate.The deviation was motivated by saturation in the summarize condition.
- D1–D3 estimator changes: D2 replaced the locked Llama-Guard-3-8B guardrail with a production-grade PII regex because Llama-Guard-3-8B was not run.The regex served as the deployed defender baseline.
- D1–D3 estimator changes: D3 replaced the locked post-guard final-leak headline metric with pre-guard generator emission, producing an estimand shift.The locked end-to-end privacy-risk estimator was not run.
- D4–D6 evaluation changes: D4 used a single-seed Qwen-2.5-VL-32B-Instruct text-only run, which the paper does not characterize as a venue-grade replication.This limits replication strength within the analyzed setup.
- D4–D6 evaluation changes: D5 used a gold-only force-context regime that violated L(QR)≈L(Q0) at 7B by −22 pp with p=0.003.The reported deviation concerns the force-context sanity relationship.
- D4–D6 evaluation changes: D6 analyzed es-LATAM per-predicate variance with only npairs=3−12 and bank-source labels.The deviation reflects small per-predicate sample sizes.
- D7 predicate-bank scale: D7 reduced the predicate bank from approximately 150 stereotype predicates per culture to 43 total, with novel sub-banks below the planned minimum.The novel sub-bank was reduced to 4/3/3 for es-LATAM, Arabic, and Hindi versus a planned ≥5.
- D8 expanded-QC sensitivity: D8 added a post-hoc seven-predicate culture-neutral QC pool rerun on QC only, using the same documents and length matching.The sensitivity run was single-seed and predicate-imbalanced.