Source-linked AI summary
Detecting Authorship in Political Texts with Inductive Stylometry
Gennadii Iakovlev, Levente Littvay
TL;DR
Political texts may carry stylistic traces from staff and other hidden contributors, but political science has paid limited attention to recovering them. This paper stress-tests frequency-based stylometry using character 3-grams, UMAP, and Burrows’ Delta across six corpora varying in length, modality, and language. The approach recovers authorship or production signals in five of six corpora, while individual speechwriters remain unresolved in scripted speech.
Problem
Political science has paid limited attention to stylistic traces left by staff and other contributors who draft or edit texts attributed to nominal political authors.
Method
The paper combines character 3-gram features with UMAP and Burrows’ Delta, applying classical frequency-based stylometry across six corpora differing in length, language, and modality.
Results
The classical stylometric toolbox recovers important authorship or production signals in five of six politically diverse corpora, including English and Hungarian texts and written and oral material.
Takeaways & Limitations
Validated stylometric signals can support analysis of staff involvement, speechwriter turnover, institutional-text authorship, and the limits of authorship proxies in political communication.
Takeaways & Limitations
The method does not identify individual speechwriters in scripted speeches, where collaborative editing, limited text per writer, and oral delivery may homogenize or obscure styles.
Abstract
from arXiv · showhide
Political texts are rarely authored by the nominal speaker alone. Tweets, speeches, reports, and official statements are drafted, edited, or harmonized by staff, yet political science has paid limited attention to the stylistic traces these hidden authors leave behind. This paper develops and stress-tests an inductive stylometric approach for recovering latent authorship structure in political communication, combining character 3-gram features with UMAP dimensionality reduction, and Burrows' Delta. We apply the approach to six corpora that vary in length (from tweets to long documents), in mode (written and oral), and in language (English and Hungarian). The approach recovers near-disjoint analyst fingerprints in formal legal prose in both languages, sorts a politician's tweets into validated subsets while uncovering additional insights, and distinguishes scripted from improvised speech. It fails, however, to resolve individual speechwriters within scripted corpora. Frequency-based stylometry is thus a powerful tool that, depending on authorial signal strength and institutional editing, can uncover authorship traces relevant to legislative studies, political communication, and policy research.
1 Introduction
The paper applies and stress-tests inductive stylometry across political texts that differ in length, modality, and language. It combines character 3-grams, UMAP, and Burrows’ Delta to recover stylistic and production signals while testing their limits.
- Motivation: Political science has largely overlooked stylometry despite political texts often being drafted or edited by staff and other contributors.The paper argues that nominal authorship may obscure the staff contributions shaping speeches, laws, court decisions, and diplomatic documents.
- Research design: The study tests classical stylometric methods across long and short texts, written and delivered speech, and English and Hungarian.The design deliberately stretches the approach across text length, modality, and a typologically divergent language.
- Findings: The approach separates basic stylistic clusters in short and long political texts, but individual speechwriter clusters do not emerge from scripted speeches.Stylometry instead sharpens distinctions between scripted and improvised delivery and other hypothesized groups.
- Method: UMAP provides an unsupervised, neighborhood-preserving visualization of stylistic similarity without using authorship labels.The method maps high-dimensional document-feature matrices into two dimensions and can emphasize local or global structure through its neighbor count.
- Method: The core analysis combines character 3-gram features projected with UMAP and Burrows’ Delta on frequent-word profiles.Character trigrams capture sub-word regularities, while Burrows’ Delta compares relative use of frequent words, especially function words.
3 Use Case 1: CRS Reports (American Law, Short-Report Series)
The controlled CRS short-report corpus isolates analyst style by holding topic, audience, register, and format approximately constant. Character 3-gram UMAP and Burrows’ Delta recover strongly separated analyst fingerprints, although specialization remains a potential confound.
- Corpus and design: The short-report series contains 167 CRS documents whose topic, audience, legal register, and format are held approximately constant.The design uses the RS format to reduce variation caused by document length and templates.
- Results: +0.47 mean silhouette width marks the study’s cleanest analyst separation, with the three most prolific analysts occupying effectively non-overlapping islands.Keith Bea, Robert Keith, and Charles Doyle form distinct regions, while four less prolific analysts occupy more diffuse but coherent neighborhoods.
- Results: Burrows’ Delta corroborates the UMAP projection because within-analyst frequent-word profiles are tighter than cross-analyst profiles.The measure provides an independent frequency-based check on the analyst clusters.
- Caveat: Analyst identity and subject specialization remain empirically entangled in the controlled slice.Several prolific analysts concentrate their reports in distinct substantive areas, so the design cannot eliminate every confound.
4 Use Case 2: CRS Reports (Full American Law Subset)
Pooling all CRS American Law formats tests whether analyst signals survive template and length variation. Analyst clusters remain strong, but format heterogeneity creates within-author splits and requires interpretation beyond visual clustering.
- Results: Analyst clusters remain clearly separated across pooled formats, although several analysts occupy multiple islands.Keith Bea, Charles Doyle, Robert Keith, and Robert Jay Dilger resolve into recognizable regions, while the interior becomes more mixed.
- Results: +0.41 analyst-label silhouette remains close to +0.47 after pooling 1,109 reports across formats.The small decline indicates that authorship signal remains stronger than the format signal in the pooled corpus.
- Validation: A topic-label control yields silhouette −0.31 versus +0.41 for analyst labels in identical coordinates, weakening subject specialization as the primary explanation.The comparison supports an authorial interpretation of the analyst structure while retaining the need for supporting validation.
- Interpretation: The pooled projection reveals format-driven within-author splits caused by differences in templates, length, boilerplate, and possibly career-era or specialization changes.Analysts who form single clusters under format control fragment into principal clusters and satellite groups when formats are mixed.
- Implications: Known analyst labels validate otherwise exploratory islands as primarily authorial regions and support future attribution of anonymous CRS and policy documents.The labeled corpus provides the bridge from exploratory clusters to authorship questions in documents without named authors.
6 Use Case 4: Trump Tweets by Device
Device-based stylometry separates Trump’s tweets, but formatting and production register account for much of the distinction, making device only a noisy authorship proxy.
- Device-specific UMAPs overlap substantially, with iPhone tweets favoring link- and hashtag-heavy campaign registers and Android tweets favoring unformatted prose.The observed pattern is more consistent with production-register differences than a clean authorship division.
- Burrows’ Delta profiles computed on chronological 50-tweet chunks make the Android and iPhone groups almost perfectly separable.The analysis aggregates short tweets into 50-tweet chunks because individual tweets are much shorter than texts for which Delta was developed.
- After promotional formatting is excluded, iPhone staff posts remain shorter and more positive, while much of their prose uses Trump’s own idiom.This residual difference follows promotional register and includes first-person attacks, signature epithets, and self-promotion.
- Device is therefore a noisy authorship proxy: Delta distinguishes the labels, but the contrast cannot be interpreted as a clean difference between authors.Leave-one-out reference profiles leave the device ranking essentially unchanged, ruling out self-inclusion as the explanation for separation.
- 0.21, 0.09, and 0.08 are the mean device silhouettes with formatting intact, without URL characters, and without URL and hashtag characters, respectively.The sharp reduction shows that formatting accounts for most of the device separation in the character 3-gram projection.
7 Use Case 5: Trump Speeches (Teleprompter vs. Off-the-Cuff)
Character 3-gram UMAP and Burrows’ Delta distinguish off-the-cuff from teleprompter-assisted speeches and capture degrees of improvisation, but do not recover individual speechwriters.
- Delivery-style separation: 23 speeches split into non-overlapping off-the-cuff and teleprompter groups in the character 3-gram UMAP projection.The groups are separated despite not being spatially distant.
- Supporting diagnostics: Sentiment is the only dictionary feature with a significant group gap, with impromptu speeches markedly more positive, while speech length does not differ.The trigram–UMAP projection separates the groups, unlike the reported t-SNE and truncated-SVD projections of TF–IDF features.
- Delivery-style separation: The projection captures a gradient of scriptedness, with teleprompter speeches farther from the scripted core containing more improvised delivery.Speeches with frequent short ad-libs occupy intermediate positions, while nearly fully scripted addresses remain at the scripted end.
- Speechwriter attribution: The scripted corpus contains no internal subclusters attributable to individual speechwriters.Collaborative editing, limited text per writer, and Trump’s delivery features may homogenize or obscure individual stylistic signals.
- Methodological qualification: UMAP coordinates provide descriptive corroboration only because they have no meaningful units and induce dependence across observations.The study therefore treats the embedding as an exploratory visualization rather than a conventional inferential test.
- Delivery-style separation: Burrows’ Delta shows a similar long tail toward the improvised profile among scripted speeches, while function words still separate delivery modes almost perfectly.The preference scores therefore corroborate the continuous pattern visible in the UMAP projection.
9 Discussion
The paper shows that classical stylometry can recover authorship or production signals across diverse political corpora, while its success depends on language, editing, delivery, and text length. It also cautions that projections require supporting validation and that authorship signals can affect interpretation of nominally authored communication.
- Five of six politically diverse corpora yielded important authorship or production signals using character 3-grams, Burrows’ Delta, and UMAP.The approach generalized across tweets and long reports, English and Hungarian, and written and oral modalities.
- Four conditions govern success: morphology, institutional editing, oral delivery, and sufficient text length.The paper contrasts independently drafted formal texts with heavily edited or orally delivered speech, where drafter-level distinctions disappear.
- UMAP clusters cannot establish authorship alone because topical, formatting, and production differences can produce apparent separation.The analysis therefore combines projections with Burrows’ Delta, labels, and discriminant-validity tests.
- Stylometric traces can complicate claims about a nominal author’s beliefs, intentions, or rhetoric when staffers or institutional hands shaped the text.The paper identifies this risk across tweets, speeches, and policy documents.
- The Orbán corpus combines adverse conditions, motivating learned authorship representations as a possible alternative where surface frequency features lose signal.The proposed direction is explicitly framed as future work rather than a demonstrated result.
B Burrows’ Delta for the CRS Corpora (Use Cases 1–2)
Burrows’ Delta supports the CRS authorship findings by showing tighter within-analyst profiles than cross-analyst profiles in both CRS corpora. The speech evidence separately shows that teleprompter coding captures heterogeneous delivery, including substantial improvisation in some coded speeches.
- Burrows’ Delta for the CRS Corpora: Within-analyst Burrows’ Delta profiles are markedly tighter than cross-analyst profiles in both the short-report and pooled CRS corpora.The measure uses distances to each analyst’s frequent-word style profile, where lower values indicate greater within-analyst consistency.
- Burrows’ Delta for the CRS Corpora: The CRS Delta evidence is consistent with character 3-gram UMAP projections and relies on frequent function words with limited topical signal.Together, these diagnostics provide complementary support for analyst-level stylistic separation.
- Trump Speeches: The character 3-gram UMAP separates most teleprompter and non-teleprompter speeches into distinct regions.Most non-teleprompter speeches occupy the upper portion, while most teleprompter speeches form a separate group.
- Trump Speeches: Teleprompter presence does not guarantee continuous reading from prepared text, because some coded speeches include brief or extended off-script passages.The West Palm Beach address contains repeated deviations and improvised stretches, while other teleprompter speeches adhere more closely to the script.
- Trump Speeches: The teleprompter category therefore includes strongly scripted speeches, speeches with short impromptu segments, and possibly largely improvised speeches.This heterogeneity qualifies interpretation of the delivery-based clustering.
D.1 Rationale and Method
The discriminant-validity analysis tests whether character 3-gram UMAP structure reflects authorship rather than surface-level confounds. In the controlled short-report corpus, several alternative factors correlate with analyst-aligned clusters, while length and sentence-level measures are generally weak.
- Rationale and Method: The analysis re-colors the same character 3-gram UMAP by six alternative factors to test surface-level explanations for clustering.The factors are sentiment, text length, mean sentence length, mean word length, passive-voice density, and coordination/subordination ratio.
- Rationale and Method: Pearson correlations between each factor and each UMAP dimension are used to assess whether systematic gradients align with authorship clusters.Low correlations and absent matching gradients make potential confounds less plausible explanations for the observed structure.
- Rationale and Method: The short-report corpus holds topic, audience, genre, and format constant, so remaining surface-factor associations reflect differences among analysts more directly.This design provides the study’s most tightly controlled setting for discriminant-validity analysis.
- Observed Pattern: Text length correlates weakly with the short-report projection (rU1 = 0.29, rU2 = −0.03), as do mean sentence length and passive-voice density.Mean sentence length has rU1 = 0.24 and rU2 = −0.19, while passive-voice density has rU1 = −0.17 and rU2 = −0.02.
- Observed Pattern: Coordination/subordination ratio, sentiment, and mean word length show stronger correlations that align with analyst clusters.The strongest reported magnitudes are rU1 = −0.63 and rU2 = 0.50 for coordination/subordination ratio, rU2 = −0.59 for sentiment, and rU2 = 0.48 for mean word length.
D.3 Use Case 2: CRS Reports (Full American Law Subset)
The full CRS American Law analysis tests whether analyst clusters are explained by topic, document scope, or surface features rather than authorship. Topic labels are heavily interspersed, while several weak length-related associations and stronger syntactic or sentiment correlations remain consistent with analyst-level variation.
- D.3 Use Case 2: CRS Reports (Full American Law Subset): Every full-corpus document shares topic area, institutional audience, and formal genre, but format and report scope remain possible confounds.The analysis considers whether differences in report length and specialization contribute to the projection.
- D.3 Use Case 2: CRS Reports (Full American Law Subset): Length, mean sentence length, and mean word length show weak associations with the full-corpus projection, including |r| ≤0.11 for the first two measures.The largest correlations involve coordination/subordination ratio, dictionary sentiment, and passive-voice density.
- D.3 Use Case 2: CRS Reports (Full American Law Subset): The strongest full-corpus correlations are rU1 = −0.50 for coordination/subordination ratio, rU2 = −0.40 for dictionary sentiment, and rU2 = −0.34 for passive-voice density.These patterns are consistent with analyst-level differences in syntactic architecture and affective framing, although sentiment requires caution because legal vocabulary coverage is limited.
- D.3 Use Case 2: CRS Reports (Full American Law Subset): The topic partition has silhouette −0.31, compared with +0.41 for author labels in the same coordinates, and is heavily interspersed.This comparison tests whether official CRS topic tags explain the analyst-aligned projection structure.
- D.3 Use Case 2: CRS Reports (Full American Law Subset): Across topics, 36 analysts contributed 942 reports while qualifying with at least six single-authored reports in each of two or more primary topic areas.The author-portability analysis fixes the author and varies the topic using character 3-gram profiles.
D.4 Use Case 3: Hungarian Ombudsman Reports
The rapporteur-level structure in Hungarian Ombudsman reports is not explained by the tested surface factors, supporting an authorship interpretation. However, two syntactic diagnostics are unreliable because they use English-language heuristics on Hungarian text.
- Discriminant validity: The topic partition is heavily interspersed in the same UMAP coordinates, with silhouette −0.31 versus +0.41 for author labels.This comparison tests whether topic structure accounts for the observed author-labelled separation.
- Discriminant validity: Dictionary sentiment shows no visible affective gradient organising the projection, making sentiment an unlikely explanation for rapporteur-level structure.The reports use a highly standardised, neutral legal-administrative register with limited affective vocabulary.
- Discriminant validity: Text length is near-zero to weakly associated with the projection, so the UMAP is not primarily sorting reports by complexity or scope.A strong length association would have suggested that report complexity, rather than authorship, drove the structure.
- Caveat: Passive-voice density and coordination/subordination are only rough proxies because their measures rely on English-language heuristics applied to Hungarian.These diagnostics should therefore be interpreted cautiously rather than as substantive syntactic evidence.
- Discriminant validity: All six tested surface factors correlate weakly with both UMAP dimensions, leaving none as a strong explanation of rapporteur-level clustering.Mean sentence length reaches |r| ≤0.24 and dictionary sentiment reaches rU1 = 0.19; text length, mean word length, passive voice, and coordination/subordination are closer to zero.
D.5 Use Case 4: Trump Tweets (Android vs. iPhone)
The tweet projection separates Android and iPhone streams, but the separation is largely associated with production register and formatting rather than a clean authorship split. Leave-one-out Burrows’ Delta checks preserve the reported device separation without resolving that interpretive limitation.
- Interpretation: Formatting accounts for most of the device separation because iPhone tweets contain links, hashtags, and promotional posts, whereas Android tweets contain composed prose.Production register covaries with length and surface form, complicating a direct authorship interpretation.
- Discriminant validity: Mean word length (rU1 = 0.52), mean sentence length (rU1 = −0.48), and text length (rU1 = −0.23, rU2 = −0.35) correlate with the device-separating axes.Sentiment, passive-voice density, and coordination/subordination show effectively zero correlations.
- Robustness check: Leave-one-out Burrows’ Delta changes mean device scores by under half a percent and leaves the device ranking and separation unchanged.The recomputation removes each tweet from its own device reference profile, ruling out self-inclusion as the source of the reported separation.
- Interpretation: The validated device contrast remains a production-register label rather than evidence of an authorial hand.The stripping analysis identifies the label with production register, so the device distinction cannot be treated as a clean authorship split.
D.6 Use Case 5: Trump Speeches (Teleprompter vs. Off-the-Cuff)
In Trump’s speeches, stylometry sharply distinguishes scripted teleprompter delivery from improvised delivery, but the separation partly reflects register differences within one speaker. Function-word Burrows’ Delta adds separation beyond the measured surface markers.
- Interpretation: The primary contrast is between scripted teleprompter and improvised ad-lib delivery by the same speaker, not between two individuals.The corpus therefore tests delivery mode and institutional versus individual voice rather than direct writer identity.
- Discriminant validity: Sentiment (rU1 = 0.55, rU2 = −0.87), mean sentence length (rU1 = −0.40, rU2 = 0.83), and mean word length (rU1 = −0.46, rU2 = 0.82) have the study’s highest correlations.Passive-voice density reaches rU1 = −0.56 and rU2 = 0.63, while coordination/subordination correlations are at most 0.52 in absolute value.
- Discriminant validity: Text length remains weakly associated with the projection (rU1 = −0.32, rU2 = 0.10), consistent with the non-significant length difference between delivery modes.The corpus therefore does not support duration as the main driver of the separation.
- Interpretation: Function-word Burrows’ Delta produces near-perfect separation between the two groups, despite function words carrying minimal affective or syntactic-complexity signal.This separation remains alongside the measured register markers, so surface register explains part but not all of the structure.
- Interpretation: Scripted text has longer sentences, more formal vocabulary, and denser passive constructions, whereas improvised speech is less formally structured.These register differences explain part of the observed contrast but not the full function-word separation.
D.7 Use Case 6: Orbán Speeches in Hungarian
The Hungarian Orbán speech corpus does not yield stable author-level clusters, while the tested factors weakly reflect speech formality and genre. Interpretation is constrained because two syntactic measures use English-language heuristics with limited validity for Hungarian morphology.
- Main finding: Frequency-based stylometry fails to recover author-level structure in the amorphous, non-clustered Orbán UMAP.The corpus has no authorship labels, so the discriminant-validity battery is the primary diagnostic.
- Discriminant validity: Mean sentence length (rU2 = −0.42), dictionary sentiment (rU2 = −0.30), and mean word length (rU1 = −0.27) are the strongest reported associations.These moderate correlations suggest weak organisation by speech formality and genre rather than authorship.
- Interpretation: The projection’s weak formality and genre pattern is consistent with differences between programmatic speeches and shorter, more formulaic ceremonial or bilateral statements.Hungarian morphology also makes mean word length sensitive to compounds and case suffixes, which may reflect topic or genre rather than authorial preference.
- Discriminant validity: Text length is near zero (rU1 = 0.05, rU2 = −0.05) despite the corpus’s wide length range, so duration does not simply organise the projection.No single tested factor dominates the layout, although register and sentence-level formality weakly shape it.
- Caveat: Passive voice and coordination/subordination are near zero, but their values have limited validity because both measures use English-language heuristics on Hungarian.Any correlations from these diagnostics should be treated cautiously rather than as substantive stylometric signals.