Source-linked AI summary
Flesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic Models
Yo Ehara
TL;DR
The paper investigates whether stable long-document FRE and FKGL scores remain invariant to lexical composition. It derives their topic-model limits through two scalar rates and audits FKGL with split-half prediction, finding strong associations but dataset- and fit-dependent incremental value. The authors restrict the interpretation to modelled lexical composition rather than human readability or causal effects.
Problem
Long-document stability of FRE and FKGL does not establish invariance to lexical composition, motivating analysis of their limits under explicit topic models.
Method
The paper derives almost-sure score limits and invariance geometry under topic models, then evaluates out-of-fold split-half FKGL prediction with masked inputs, controls, and lexical baselines.
Results
r = 0.779 on Brown and r = 0.884 on the BNC show strong cross-half FKGL prediction; Brown’s incremental ΔR^2 = 0.002 is indistinguishable from zero, while the BNC increment is fit-dependent.
Takeaways & Limitations
Under the model, long-text FKGL is determined by document-level lexical composition through two scalar rates, while the experiments do not establish a human-readability or causal interpretation.
Takeaways & Limitations
The theory’s O_p(N^-1/2) rate assumes conditional independence, and the empirical scope is limited to two corpora, one language, and fit-dependent topic-model audits.
Abstract
from arXiv · showhide
Flesch Reading Ease (FRE) and the Flesch-Kincaid Grade Level (FKGL) are widely used readability scores for English computed from the same two document statistics, yet their stability on long documents need not imply invariance to lexical composition. Surprisingly, under a topic model with an explicit sentence-boundary token, both scores converge almost surely to deterministic functions of the document topic distribution through just two scalar rates: in the long-text limit, all score variation is mediated by topical composition rather than any residual readability signal. The theory covers both formulae, while the experiments evaluate FKGL. In a fixed admixture with rank[1, q, s] = 3, fibres through interior topic vectors are locally (K-3)-dimensional, whereas regular iso-score level sets are locally (K-2)-dimensional and curved. In out-of-fold evaluation on two balanced corpora, Brown and the written BNC, a topic vector inferred from one document half's content words predicts the other half's FKGL at r = 0.779 and 0.884, respectively. On Brown, adding the topic prediction to genre and mean content-word syllable count yields $ΔR^2$ = 0.002, with a confidence interval spanning zero; on the BNC, the corresponding split-half increment is 0.024, positive in four of five K = 100 fits (median 0.021). Because inferred topics may also absorb genre, register, and style, we do not interpret these results as evidence about human readability or causal effects.
1 Introduction
The paper asks whether long-document stability of FRE and FKGL also implies invariance to lexical composition. Under topic models with an explicit sentence-boundary token, both scores converge to deterministic functions of document topic proportions, with experiments showing strong split-half FKGL prediction.
- Motivation: FRE and FKGL use the same two statistics: words per sentence and syllables per word.The theory treats both formulae through shared rate-based quantities, while the experiments evaluate FKGL.
- Theory: Long-text FKGL converges almost surely to a closed-form function Φ(θ) of the document topic vector under conditionally i.i.d. topic-model tokens.The convergence has O_p(N^-1/2) error and an explicit asymptotic variance; FRE follows by coefficient substitution.
- Theory: The limiting score depends on two scalar rates: the sentence-boundary rate and expected syllable contribution per token.This follows because words per sentence is the reciprocal of the boundary-token rate and syllable count is determined by word type.
- Scope: The findings concern modelled lexical composition, not human readability or causal effects, because inferred topics can absorb genre, register, and style.The corpora contain no human judgements.
- Geometry: For rank[1, q, s] = 3, interior fibres are locally (K - 3)-dimensional, while regular iso-score sets are curved and locally (K - 2)-dimensional.Along a two-topic mixing path, the score is a ratio of polynomials of degree at most two rather than an average of endpoint scores.
- Experiments: r = 0.779 on Brown and r = 0.884 on the written BNC quantify out-of-fold prediction of one half’s FKGL from the other half’s inferred topic vector.Adding topic prediction to genre and mean content-word syllable count gives ΔR^2 = 0.002 on Brown, with a confidence interval spanning zero, and 0.024 on the BNC.
2 Related Work
Related work frames readability formulas as simple sentence-length and word-difficulty measures, while this paper analyzes their structural sensitivity under an explicit topic model. The rate-based formulation connects both terms to token-level rates and clarifies the topic-model assumptions needed for the theory.
- Readability formulae and their critique: Classical readability formulas combine a sentence-length term with a word-difficulty term, but have documented limitations including gameability and weak correlation with human judgements.Modern assessment instead uses supervised models with richer features or contextual encodings.
- Rate-based readings of FKGL: Both formula terms can be written as reciprocals of token rates when each sentence ends with an explicit boundary symbol.The resulting rates include sentences per word and words per syllable, whose product gives syllables per sentence.
- Rate-based readings of FKGL: For observed documents, the rate factor is bounded because preprocessing ensures 0 < S ≤ W ≤ Y.This yields 0 < p_wℓ, p_sw ≤ 1 and bounds the coefficient-dependent factor M.
- Topic models: LDA, ProdLDA, and ETM share a document-level topic vector that induces conditionally i.i.d. token distributions, the property required by the theory.NMF can serve as a normalized admixture control, whereas truncated SVD generally cannot.
3 FKGL under Topic Models
Under an explicit sentence-boundary topic model, long-document FKGL converges to a deterministic function of topic composition through two scalar rates. The resulting geometry shows which topic changes preserve the score, while experiments test how much FKGL is recoverable from inferred topics.
- Setup: A token-exchangeable topic model emits words and sentence boundaries from a document topic vector θ, with syllable counts determined by word type.The boundary token supplies sentence counts, while bounded token-level averages support the asymptotic argument.
- Long-text limit: FKGL converges almost surely to Φ(θ), because the boundary-token rate e(θ) and expected syllable contribution σ(θ) determine the limiting formula.The proof applies the strong law of large numbers and the continuous mapping theorem; the same framework covers dependent stationary ergodic tokens.
- Two-rate factorisation: The limit factors as Φ = g ◦ (e, σ), so topic vectors sharing both rates receive the same asymptotic score.For admixture models, e(θ) = q^Tθ and σ(θ) = s^Tθ; ProdLDA retains the outer form but makes these rates nonlinear in θ.
- Invariance geometry: For rank[1, q, s] = 3, interior fixed-rate fibres are locally (K −3)-dimensional, whereas regular iso-score level sets are locally (K −2)-dimensional and curved.The extra level-set direction varies with θ, and the result depends on the stated rank and interiority conditions.
- Mixing paths: Along a two-topic mixing path, Φ is generally a ratio of polynomials of degree at most two and need not equal the endpoint-score average.Only the syllables-per-sentence factor is linear-fractional; the complete score is not generally affine in the mixing weight.
- Rate of convergence: The asymptotic error is O_p(N^-1/2), but convergence slows as the boundary rate e approaches zero because sentence counts become rarer.The asymptotic variance follows from a multivariate CLT and delta method, with a constant that grows as e → 0.
4 Experiments
The experiments test whether FKGL can be recovered from topic vectors inferred out of fold, using masked inputs and disjoint document halves. Topic predictions correlate strongly with held-out FKGL, but incremental value beyond genre and mean content-word syllable count differs between Brown and the BNC.
- 4.1 Setup: Out-of-fold evaluation uses five-fold model fitting, disjoint document halves, masked inference inputs, genre controls, and a mean content-word syllable baseline.The corpora are Brown and the written BNC; the BNC contains substantially longer documents than Brown.
- 4.4 How much of this is trivially lexical?: ΔR2 = 0.002 [−0.003, 0.007] on Brown when Φ(θ̂) is added to genre and BC, so the confidence interval spans zero.The topic pipeline matches the simpler genre-plus-BC baseline statistically in this split-half analysis.
- 4.4 How much of this is trivially lexical?: ΔR2 = 0.024 [0.018, 0.031] on the BNC for adding Φ(θ̂) to genre and BC in the split-half analysis.The increment is positive in four of five K = 100 fits, with median 0.021.
- 4.3 Out-of-fold agreement: A topic vector inferred from one half’s content words predicts the other half’s FKGL at r = 0.779 on Brown and 0.884 on the BNC.These are cross-half correlations using content-masked inference.
- 4.7 Convergence rate and plug-in bias: Agreement rises with segment length, from 0.556 to 0.858 on Brown and from 0.587 to 0.930 on the BNC.The passage describes this as descriptive rather than a direct verification of the theorem.
- 4.7 Convergence rate and plug-in bias: The asymptotic standard deviation is within ∼3% of empirical variability at N = 1,600 but underestimates variability for N ≤ 200.The closed-form variance is therefore treated as a long-text approximation, not a finite-sample guarantee.
5 Discussion
The discussion frames FKGL as strongly associated with document-level lexical distributions compressed by topic models, while emphasizing that the association is not itself evidence about readability or causality. It also notes that FKGL can be changed through lexical-distribution shifts without establishing simplification.
- What the experiments license: FKGL is strongly predictable out of fold from the document-level lexical distribution compressed by a topic model, and the association is not exhausted by coarse genre labels.On Brown, the incremental signal is largely exhausted by genre plus BC; on the BNC, a positive increment appears in four of five K = 100 fits.
- Consequences: A generator can lower FKGL by shifting its lexical distribution toward smaller Φ, or move along fibre directions with FKGL constant.FKGL alone cannot establish whether such shifts constitute simplification, including for FKGL-band filtering.
6 Conclusion
The theory makes FKGL a deterministic function of topic composition in long texts, and split-half experiments show strong cross-half prediction in Brown and the BNC.
- Conclusion: FKGL converges almost surely to a closed-form function of the topic vector, with Op(N^-1/2) error.For rank[1, q, s] = 3, interior fibres are locally (K-3)-dimensional, while regular level sets are locally (K-2)-dimensional and curved.
- Conclusion: r = 0.884 in the BNC and r = 0.779 in Brown for predicting one half’s FKGL from the other half’s content-word topic vector.
- Conclusion: The one-feature baseline matches the topic prediction on Brown, while the BNC increment is positive in four of five K=100 fits.
Limitations
The theoretical rates rely on assumptions that real text violates, and the latent topic vector conflates subject matter with genre, register, style, and typographic convention.
- Theory: The conditional i.i.d. assumption is false of real text; under stationarity and ergodicity, only almost-sure convergence is established.The Op(N^-1/2) rate and asymptotic variance result rely on independence and understate variance under sentence-length autocorrelation.
- Interpretation: The topic vector is not semantics because a bag-of-words latent absorbs genre, register, style, typographic convention, and subject matter.On Brown, mean content-word syllable length plus genre matches the recoverable signal; the data do not separate subject matter from style.
- Method: Adding an explicit boundary token supplies one of FKGL’s two sufficient statistics under the bag-of-words model.Masked variants alter inference inputs, while q_k remains estimated from training documents containing the boundary token.
- Empirical scope: The empirical scope is limited to two English corpora and LDA, with model-fit and fold-partition sensitivity remaining relevant.Across six fits, the whole-document increment was positive, while the content-masked split-half increment was negative in one K=100 fit.
Ethical Considerations
The paper reports no direct ethical risk in its mathematical analysis and corpus experiments, but warns that FKGL-based downstream filters can disadvantage texts with particular lexical profiles.
- Ethical Considerations: FKGL-based automatic rewards or filters can systematically disadvantage texts with particular lexical profiles.Those profiles correlate with subject matter and genre, regardless of qualities the data cannot measure.
- Ethical Considerations: The analysis uses two licensed corpora of published English: Brown and the British National Corpus.
- Ethical Considerations: The authors report no direct ethical risk in the mathematical analysis and experiments themselves.
D Convergence rate and plug-in bias
Simulation results validate the asymptotic variance approximation for sufficiently long texts, while short texts exhibit heavy-tailed instability; posterior plug-in corrections barely change correlations.
- Convergence rate: The observed-to-predicted standard-deviation ratio approaches 1 as N increases: 1.18 at 400 tokens, 1.08 at 800, 1.03 at 1,600, and 1.015 at 3,200.
- Convergence rate: Below N ≈200, the delta method substantially underestimates variability because the sentence count is small and W_N/S_N is heavy-tailed.
- Convergence rate: A 0.5-grade standard deviation requires approximately 5,700 tokens when τ is averaged across 60 documents.
- Plug-in bias: Posterior-predictive plug-in correlations are nearly unchanged from plug-in correlations: 0.861/0.807/0.805 versus 0.861/0.809/0.805 for LDA/ProdLDA/ETM.
- Reconstruction controls: Rank-50 SVD and NMF achieve r = 0.945 and 0.918 on the unablated Brown task, but fall to 0.112 and 0.023 when boundary and stopword inputs are deleted.The unablated comparison measures reconstruction fidelity rather than topic inference.
F Confidence intervals, error and calibration
The Brown out-of-fold calibration compares topic-predicted with measured FKGL and reports prediction dispersion, error, and sensitivity to target representation.
- Figure 4 plots Brown LDA topic-predicted FKGL against measured FKGL using identity and least-squares-fit reference lines.
- Table 3 reports out-of-fold, uncalibrated FKGL predictions, with slopes, RMSE in grade levels, and R2 defined from mean-squared error.
- The neural Brown models are under-dispersed, with slopes 1.79 and 1.25, consistent with shrinkage from amortised inference.
- Replacing vocabulary-consistent FKGL with raw-surface-form FKGL changes Brown correlations to 0.866, 0.811, and 0.810 for LDA, ProdLDA, and ETM.
- Table 4 presents simplex-vertex topic examples rather than observed-document FKGL values.
H Combined controls and lexical baseline
The combined-control analyses compare topic predictions with lexical and genre baselines, including a word-only supervised rate probe and BNC split-half results.
- The word-only probe maps topic vectors to sentence-length and syllable rates using ridge regression fitted on outer training folds.
- The BNC whole-document word-only probe reaches r = 0.911, compared with 0.908 for boundary-augmented LDA.
- Under content masking and split-half evaluation, component correlations are 0.552 versus 0.812 for Brown sentence-length rates and 0.762 versus 0.900 for BNC rates.
- The BNC split-half baseline audit reports Δr(Φ − BC) = 0.039 [0.031, 0.048] and ΔR2 = 0.024 [0.018, 0.031].
I Robustness
Robustness analyses examine seed sensitivity, genre transfer, topic-count protocols, tokenisation, and corpus and computational implementation boundaries.
- Across seeds, whole-corpus LDA correlations are 0.870, 0.869, and 0.884, while per-topic quantities vary substantially.
- Observed-document Φ ranges remain similar across seeds, spanning [4.47, 14.07], [4.66, 14.18], and [4.04, 14.46].
- Deleting Brown’s five formula-placeholder types leaves corpus FKGL unchanged to two decimals and changes LDA correlation from 0.870 to 0.871.
- Leave-one-genre-out LDA reaches r = 0.642 from informative to imaginative prose and 0.425 in the reverse direction.
- Table 9 uses an in-sample protocol, so its Pearson and Spearman results are not comparable in level to Table 1.
- The BNC length-curve figure uses segment-level bootstrap 95% CIs that are optimistic because segments share documents.