Source-linked AI summary

Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches

Marco Simnacher, Georg Keilbar, Benjamin König, Christoph Lippert, Sonja Greven

arXiv:2609.00946v1stat.MLcs.AIcs.LGmath.STstat.ME

TL;DR

Existing conditional independence tests have limited applicability to high-dimensional textual data, although they are useful for testing whether LLM outputs add protected-attribute information beyond their source text. The paper proposes embedded CITs with sufficiency-based validity conditions and evaluates them through simulations and German Parliament summaries, finding information about speakers’ faction and gender beyond the speeches.

  • Problem

    Existing conditional independence tests have limited applicability to high-dimensional multimodal data such as text, despite the need to test information in LLM outputs beyond source texts.

  • Method

    Embedded CITs map output and source texts to representations and apply an existing CIT, with sufficient or mean-sufficient source embeddings transferring validity to the textual hypothesis.

  • Results

    For both summarizers and all considered embedding-map combinations, German Parliament summaries contain information about speakers’ faction and gender beyond the source speeches.

  • Takeaways & Limitations

    eCITs provide statistical guarantees for bias claims about LLM outputs on a specified task and dataset when the representation sufficiency conditions hold.

  • Takeaways & Limitations

    Validity depends on the sufficiency of the source-text embedding under the given data-generating process, and marginalizing over learned embedding parameters need not preserve conditional independence.

Abstract

from arXiv · show

Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. However, we show that such tests are of interest for large language model (LLM) outputs, where we test whether an output $X$ generated from a source text $Z$ carries information about an attribute $Y$ beyond $Z$ itself. For this purpose, we propose embedded CITs (eCITs), which embed $X$ and $Z$ and apply an existing CIT to the resulting representations and to $Y$. We show that, provided the embedding of $Z$ is sufficient, i.e. retains the information $Z$ carries about either $Y$ or the representation of $X$, the null hypothesis transfers from $X$ and $Z$ to their representations, so that a CIT valid for the embedded hypothesis is valid for the original one. We further give conditions for equivalence of the two hypotheses, and show that sufficiency weakens to mean sufficiency when the embedded test targets conditional mean independence. We propose a semi-synthetic simulation design to assess type I error (T1E) control and power of the eCITs for given embedding maps on a specific dataset and task, and use it to evaluate them on our application. Applying the eCITs to German Parliament speeches, we find for all combinations of embedding maps considered that the summaries of two LLMs contain information about the speaker's faction and gender beyond the speech they were generated from.

1 Introduction

The paper frames LLM bias as conditional dependence between generated text and a protected attribute beyond the source text, then proposes eCITs to test this for textual inputs. It establishes embedding sufficiency conditions for validity and finds calibrated tests or inflated type I error depending on the embedding, while detecting faction and gender information in German Parliament summaries.

  • Motivation: Conditional independence tests assess whether X and Y remain independent given Z while controlling type I error and retaining power against dependence.The null is H0: X ⊥⊥ Y | Z, and the alternative is conditional dependence.
  • Motivation: For LLM evaluation, bias is defined as additional information in output X about protected attribute Y beyond source text Z.This definition is reference-free and applies to outputs generated from source texts, including summaries of parliament speeches.
  • Problem: Existing CITs are difficult to apply to textual X and Z because their structure and high dimensionality obstruct required conditional-distribution, kernel, partition, or rate-condition constructions.The resulting tests may need to condition on feature representations rather than the original source text, with consequences for validity and power.
  • Method: Embedded CITs map X and Z to feature representations and apply an existing CIT to test conditional dependence between represented X and Y given represented Z.The representation of Z must retain relevant information about Y or represented X; dropped information can act as an omitted conditioning variable and inflate type I error.
  • Theory and contributions: The paper characterizes sufficient embeddings of Z that transfer the textual null to the embedded null, gives equivalence conditions, and weakens sufficiency to mean sufficiency for conditional mean independence.It also adapts the projected covariance measure to categorical Y using a precision-weighted projection direction shown to be optimal.
  • Evaluation and application: On speech-summary pairs, eCITs with sufficient Z embeddings are calibrated across considered sample sizes, whereas insufficient embeddings can produce negligible to steeply growing type I error inflation.In the German Parliament application, all considered embedding combinations find that two LLMs’ summaries contain faction and gender information beyond the speeches.

2 Conditional (In)dependence for LLM-Generated Summaries

The paper defines summary bias through the conditional-independence relation between an LLM output, a protected attribute, and the source text. It motivates this criterion through training-data dependence and demonstrates PCM behavior under null and alternative constructions.

  • Operational definition: The tested bias hypothesis is H0: Y ⊥⊥ X | Z versus H1: Y̸ ⊥⊥ X | Z, where X is generated from source text Z and Y is protected.Under the null, the output adds no protected-attribute information beyond the source; under the alternative, it adds information absent from the source.
  • Simulation design: The simulation distinguishes the LLM training corpus from the CIT test dataset and generates outputs using a fixed task prompt applied to each source text.Out-of-sample predictions represent the null, while in-sample predictions represent the alternative in the described construction.
  • Null and alternative: If the LLM training corpus is independent of the test data, generated output is conditionally independent of Y given Z because it depends on Z, training data, and independent sampling noise.This provides a null construction for the testing framework.
  • Null and alternative: Memorization can induce conditional dependence when training data include a source text or other texts by the same speaker carrying information about the protected attribute.The paper considers this mechanism plausible for models pretrained on web-scale German corpora containing Bundestag protocols.
  • Simulation results: PCM rejection rates are within one Monte Carlo standard error of the nominal 0.05 level under the null, while alternative power varies with k in k-nearest-neighbor models.Smaller k gives each observation greater influence on predictions, producing larger conditional dependence and higher rejection rates in the example.

3 Embedded Conditional Independence Tests

eCITs embed text inputs and apply a conditional independence test to the resulting representations, with validity depending on assumptions about the embeddings and their parameters. The theory characterizes sufficient embeddings, validity transfer, equivalence conditions, and a weaker mean-sufficiency result for conditional mean independence.

  • 3 Embedded Conditional Independence Tests: eCITs map X and Z to feature representations and apply an existing CIT to the embedded X, Y, and embedded Z.The test may include randomized or deterministic CIT procedures through the composition of the embedding map and the underlying test.
  • 3 Embedded Conditional Independence Tests: The X embedding primarily affects power, whereas the Z embedding affects both type I error control and power.The distinction holds when the X-embedding procedure satisfies the stated learning conditions, such as sample splitting or external training.
  • 3.1 Sufficient Embedding Maps: A sufficient Z embedding retains all information in Z about either Y or the embedded X, allowing the text-level null to imply the embedded null.The embedding parameters must also satisfy an assumption preventing them from introducing the relevant conditional dependence.
  • 3.1 Sufficient Embedding Maps: If the embedded CIT is valid for its null class, the resulting eCIT is valid for the text-level null at the same significance level.This validity transfer is stated for general estimated Z-embedding parameters, including cases where their conditional sample laws are not product measures.
  • 3.1 Sufficient Embedding Maps: Equivalence of text and embedded hypotheses prevents alternatives from becoming embedded nulls in principle, but does not guarantee equal power for a specific eCIT.Embedding can reduce effect sizes or project onto alternatives that the chosen CIT cannot detect.
  • 3.2 Conditional Mean Independence: For conditional mean independence, sufficiency only requires the Z embedding to preserve the conditional mean of Y rather than its entire conditional law.In practice, the residual information omitted by the embedding should be negligible relative to n^-1/2; the paper proposes simulation-based assessment for this condition.

4 Application

The application evaluates eCIT calibration and power under semi-synthetic data-generating mechanisms before testing German Parliament summaries. Results show that embedding sufficiency strongly affects calibration and power, while both LLMs’ summaries contain information about speaker gender and faction beyond the source speech.

  • Simulation results: The oracle eCIT controls the nominal level and rejects in every replication under the alternative at n = 1,000.It provides a reference for attributing deviations in other eCITs to embedding maps or the PCM.
  • Simulation results: Under the Jina null, all eCITs remain within or near two MC standard errors of α = 0.05, with a maximum rejection rate around 13% at n = 10,000.This can hold even when ZJina is omitted if the lost information remains negligible at √n scale.
  • Simulation results: Under the Jina alternative, power depends on how much conditional signal the X and Z embeddings retain, reaching 55–97% for binary and 56–90% for multinomial Y at n = 5,000 in arms containing the generating embeddings.Adding embeddings to X can recover power, while multinomial outcomes are less powerful at the largest sample size.
  • Simulation results: Under the Qwen DGM, sufficient Z embeddings keep rejection rates at most 12% for binary and 11% for multinomial Y while reaching at least 90% power at n = 10,000.Omitting ZQwen can instead produce severe T1E inflation, such as 86% at n = 10,000 for one binary arm.
  • Simulation results: For continuous Y, sufficient diluted arms are inflated at n = 1,000 but calibrated at n = 10,000, whereas arms omitting the generating embedding remain inflated.The pattern is consistent with slower nuisance-regression convergence and omitted information acting simultaneously.
  • German Parliament application: In the application, all sixteen tests reject at α = 0.05, and the summaries of both LLMs contain information about speaker faction and gender beyond the source speech.The smallest statistic is T = 3.5 with p = 2.2 × 10^-4; larger Qwen statistics are not interpreted as a summarizer comparison.

5 Conclusion

The paper establishes eCITs as a way to test bias in LLM outputs with statistical guarantees, while emphasizing that validity depends on embedding sufficiency. It applies the method to German Parliament speeches and finds evidence of faction- and gender-related information in summaries beyond their source speeches.

  • The paper characterizes sufficient embeddings that transfer validity from the embedded hypothesis to the textual hypothesis, with mean sufficiency sufficient for conditional mean independence.
  • The proposed simulation design assesses T1E control and power for the dataset and embedding maps before application, distinguishing calibrated from inflated regimes.
  • On German Parliament speeches, eCITs reject the null for both the speaker’s faction and gender when testing two LLMs’ summaries.The authors note that four alternative explanations cannot be excluded, although calibration and consistent test-statistic changes argue against any single explanation accounting for the rejections alone.
  • Exact (mean) sufficiency remains the main open challenge because pretrained or learned embeddings may retain residual relevant information that affects T1E at larger sample sizes.
  • eCITs make conditional independence testing applicable to multimodal data and provide guarantees for bias claims about LLM outputs on specified tasks and datasets.

Use of AI Tools

The authors state that LLMs assisted with manuscript drafting and language editing, while the analyzed LLMs are the study’s objects rather than manuscript contributors.

  • The authors used LLMs to assist with drafting and language editing during manuscript preparation.
  • The LLMs analyzed in the study are described in Appendix B and serve as the objects of study rather than part of the manuscript.

Appendix A. Proofs

The appendix proves the paper’s conditional-independence and conditional-mean-independence lifting results using disintegration, monotone-class arguments, and i.i.d. sample-level reductions. These lemmas support the theoretical transfer from textual hypotheses to embedded hypotheses.

  • The proofs reduce conditional independence statements to regular conditional distributions and measurable conditional-expectation representations.
  • Monotone-class arguments reduce measurability and independence claims to countable generating algebras and extend them to the relevant sigma-fields.
  • Lemma 19 lifts conditional independence between individual variables to their i.i.d. samples and proves the converse by decomposition, redundancy, and contraction.
  • The mean-version lifting lemma establishes the analogous equivalence for conditional expectations under integrability and a Borel-valued outcome assumption.
  • The theorem proof adjoins estimated embedding parameters, replaces the source text with its embedding, and derives the embedded null using conditional-independence rules.
  • Under the sufficiency conditions, the original textual null implies the embedded null, after which a valid CIT yields a valid eCIT.

B.1 Source Texts, LLMs, and Summary Generation

The appendix cleans German Parliament speeches, filters and merges them into source texts, and generates reproducible summaries with specified LLM prompts and decoding constraints.

  • Speeches are cleaned by removing parenthetical audience reactions, procedural phrases, dashed text, extraneous whitespace, and inconsistent quotation marks.
  • Speeches are retained only with at least 50 words and two sentences, after excluding texts containing procedural indicators.
  • Consecutive speeches from each speaker are merged before entering the CIT.
  • Gemma 4 12B generates 3–4-sentence summaries with a 256-token maximum and greedy decoding for reproducibility; Qwen 3.6 27B is additionally used in the real-world application.
  • The summaries are generated from each source text using a task-constant prompt focused on concise, objective reporting of main arguments and political positions.

B.2 Embedding Maps

The study uses four multilingual pretrained embedding models plus learned German representations, with selected dimensionality reductions and normalization. Simulations combine these maps in oracle and diluted configurations for speeches and summaries.

  • Pretrained embeddings: Four multilingual models embed speeches and summaries: BGE-M3, Jina, Llama Nemotron, and Qwen3-Embedding-8B.BGE-M3 retains 1024 dimensions; Jina, Llama, and Qwen representations are truncated to 32 dimensions using Matryoshka truncation and then normalized and standardized.
  • Learned embeddings: A learned representation uses ModernGBERT hidden states trained to predict faction or gender from raw speeches or summaries.The learned 768-dimensional test representations are used alone or concatenated with pretrained embeddings in the application.
  • Simulation configurations: The simulation generates outcomes from one 32-dimensional Gemma-summary embedding and tests oracle or diluted concatenations of speech and summary maps.Diluted configurations prepend the generating embedder to additional pretrained and learned representations.
  • Outcome generation: The continuous-outcome simulation sets r^2 = 0.4, so the speech score explains 40% of Y variance under the null.Binary and multinomial noise is determined by their link functions, so outcome types do not share a common signal scale.

C.1 Jina Embedding Data Generating Mechanism

The Jina data-generating mechanism evaluates eCIT rejection rates and p-value calibration under null and alternative settings across embedding configurations. The figures include oracle, diluted, and Jina-omitting combinations.

  • Figure 3: Figure 3 reports eCIT rejection rates for binary, multinomial, and continuous outcomes generated from Jina embeddings.
  • Embedding configurations: The Jina experiment compares oracle embeddings with diluted combinations that add learned, BGE, Qwen, and Llama embeddings.It also includes corresponding combinations that exclude the Jina embedding.
  • Figure 4: Figure 4 compares observed p-values with uniform-distribution theoretical quantiles using QQ-plots for specified Jina-based eCIT embeddings.The diagonal reference is the theoretical uniform-quantile line.

C.2 Qwen Embedding Data Generating Mechanism

The Qwen data-generating mechanism evaluates rejection rates and p-value calibration for eCITs using Qwen-based embedding configurations. The accompanying methodology defines the categorical regression and calibration procedure.

  • Figure 5: Figure 5 reports eCIT rejection rates for binary, multinomial, and continuous outcomes generated from Qwen embeddings.The experiment compares Qwen oracle, diluted, and Qwen-omitting configurations.
  • Figure 6: Figure 6 uses QQ-plots to compare observed p-values with uniform theoretical quantiles when Y is generated under the null from the Qwen speech embedding.The diagonal y = x line is the reference for uniform p-values.
  • Calibration procedure: For categorical outcomes, the PCM fits multinomial logistic regressions for Y and least-squares regressions for fitted probabilities and fitted f.
  • Categorical PCM: The categorical formulation represents the conditional distribution of Y through reduced class indicators and fitted probability vectors.The PCM estimates g, mY, and h from these representations before forming the test statistic.
  • Regularization: The ridge penalty stabilizes inversion of the estimated conditional covariance when fitted class probabilities approach 0 or 1.The penalty is selected on a split training sample over a grid from 10^-6 to 1.
  • Decision rule: The test rejects at level α when T > z_1−α after nuisance estimation on the evaluation split.

D.2 Validity

The categorical PCM reduces the embedded CI problem to a conditional mean formulation while retaining the CI null, under stated nuisance-estimation and covariance conditions. Uniform asymptotic level is not claimed for this variant.

  • Test construction: The PCM statistic averages scalar inner-product projections of vector residuals, which remain conditionally mean-zero under the null given the training data.This permits the asymptotic argument of Lundborg et al. to extend with vector-valued nuisance quantities.
  • Assumptions: The validity assumptions require E1E2 = oP(n^-1/2), a bounded normalized moment condition, and λ_min(Σ(X,Z)) ≥ c0 almost surely.
  • Categorical null: For categorical Y, conditional mean independence of the reduced class-indicator vector is equivalent to conditional independence, so the PCM tests the CI hypothesis itself.The equivalence holds for the i.i.d. laws considered in Corollary 15.
  • Scope: The paper does not claim uniform asymptotic level over a nonparametric class for this variant.Null calibration is assessed instead by simulation in Subsection 4.2 and Appendix C.

D.3 The Precision-Weighted Direction

The precision-weighted direction maximizes the population signal-to-noise ratio and is characterized by the inverse covariance-weighted signal. Under positive-definiteness and integrability conditions, this optimizer is unique up to positive scaling.

  • Motivation: The weighting identifies the direction maximizing the population signal-to-noise ratio, reducing at K = 1 to f ∝ h/v.Here v = Var(˜Y | X, Z).
  • Characterization: Under the proposition’s conditions, the supremum is attained if and only if f = κΣ^-1h almost surely for some κ > 0.The conditions require Σ(X, Z) to be almost surely positive definite and 0 < E[h⊤Σ^-1h] < ∞.
  • Proof: The proof transforms the numerator and denominator using u = Σ^-1/2h and w = Σ^1/2f, then applies pointwise and L2 Cauchy–Schwarz inequalities.Equality requires w to be a positive scalar multiple of u, yielding the stated optimizer.
  • Implementation: The regularization constant stabilizes inversion of the fitted covariance when probabilities approach the boundary, without a sign correction.Because the regularized matrix is positive definite, the constructed inner products are nonnegative.

D.4 Nuisance Estimation on the Evaluation Half

Unpenalized nuisance fits on the evaluation half create exact finite-sample orthogonality, removing one regression contribution from the numerator. The remaining bias depends on nuisance-estimation errors, while high-dimensional embeddings and misspecification limit the available guarantee.

  • Orthogonality: Unpenalized multinomial-logistic and least-squares fits make both nuisance residuals orthogonal to the design matrix W.The corresponding score and normal equations give W⊤R̂Y = 0 and W⊤R̂f = 0.
  • Numerator decomposition: This orthogonality eliminates the regression of f̂ on Z from the numerator, leaving a separate bias term under the null.The conditional-mean-zero component vanishes because f̂ does not depend on Y on the evaluation half.
  • Error control: The bias bound is negligible only when the two nuisance errors are jointly small relative to n2^-1/2.The bound has order √n2(E1E2)^1/2.
  • Limitations: With embedding dimensions p = 768–1888 and evaluation samples of roughly 25,000–27,000, the sufficient rate condition is unavailable, and misspecification remains a serious concern.The simulations suggest the bound may be conservative, but they do not separate misspecification from mean insufficiency.
Loading 2609.00946v1…