Source-linked AI summary
Opinionated, Hesitant and Stressed: Three Studies of How Politicians Speak in Four Slavic Parliaments
Ivan Porupski, Nikola Ljubešić
TL;DR
Large-scale multilingual evidence on parliamentary speech remains limited across acoustics, hesitation, and stress. This paper analyzes ParlaSpeech 3.0 across four Slavic languages and finds shared cross-linguistic patterns alongside language- and domain-specific variation.
Problem
The paper addresses limited large-scale, multilingual evidence on how politicians acoustically realize sentiment, produce filled pauses, and distribute variable stress.
Method
The paper conducts three studies using ParlaSpeech 3.0, combining multilingual parliamentary speech, aligned transcripts, sentiment and acoustic analysis, filled-pause modelling, and Croatian stress annotations.
Results
Across four parliaments, negative speech is higher, louder, and faster; hesitation predictors are broadly shared except for gender differences; Croatian stress preferences cohere centrally but fragment at the periphery.
Takeaways & Limitations
Parliamentary speech is a productive domain for comparative research in which language-level variation should remain part of the signal.
Takeaways & Limitations
The studies do not model procedural heterogeneity such as the balance between scripted delivery and extemporaneous response, and the explanations for gender differences remain unresolved.
Abstract
from arXiv · showhide
We present three large-scale studies of spoken parliamentary speech across four Slavic languages (Croatian, Czech, Polish, Serbian), drawing on over 6,000 hours from the ParlaSpeech 3.0 corpus. The first study examines how utterance-level sentiment shapes acoustic realisation: negative speech is consistently produced with higher pitch, greater intensity, and faster rate across all four parliaments, with a secondary arousal-driven upturn at the most positive extreme. The second study models filled pause frequency using negative binomial GEE, finding that speech rate, age, and sentiment are robust cross-lingual predictors, while gender effects reverse between South Slavic (men produce fewer filled pauses) and West Slavic parliaments (no gender difference) - a pattern invisible to single-language designs. The third study investigates primary stress variation in Croatian, showing that speaker-level preferences for early versus late stress cohere across verbs, adjectives, and nouns but decouple for adverbs and proper nouns. We conclude with a research agenda spanning corpus phonetics, disfluency modelling, and political rhetoric.
1 INTRODUCTION
The paper uses multilingual parliamentary speech to study how politicians express sentiment, hesitation, and lexical stress. ParlaSpeech 3.0 makes these questions systematically analyzable across Croatian, Czech, Polish, and Serbian.
- Three studies ask whether sentiment affects acoustic delivery, what predicts filled pauses, and whether stress preferences remain coherent across lexical categories.
- Over 6,000 hours of Croatian, Czech, Polish, and Serbian parliamentary speech are collected in ParlaSpeech with aligned official transcripts.
- ParlaSpeech 3.0 adds speaker metadata, syntactic annotations, sentiment predictions, filled-pause detection, and stress alignments for Croatian and Serbian.
- These resources support cross-referencing sentiment with acoustics, filled pauses with syntax, and stress patterns with speaker demographics and part of speech.
- The paper presents the three investigations before outlining future research directions opened by them.
2 SENTIMENT AND ACOUSTICS
Across four parliaments, negative sentiment is associated with higher, louder, and faster speech, while the strongest positive sentiment produces a secondary upturn. Valence dominates overall, but arousal-related curvature contributes across languages and features.
- 96% of feature–language combinations showed significant sentiment-extreme differences, while monotonic trends were significant in 12/12 tests.
- Sentiment-related pitch and intensity decreased from negative toward neutral-positive speech before turning upward at the most positive extreme.
- The inflection point occurred at sentiment value 3.5, the neutral-positive boundary, and was fairly consistent across features and languages.
- 42% of tests supported secondary arousal-like curvature, with speech rate showing the strongest support at 63%.
- Negative speech was produced with higher pitch, greater intensity, and faster rate across all four parliaments.
- Croatian mainly followed linear valence, Polish showed the strongest arousal modulation, and Czech and Serbian concentrated arousal signatures in speech rate.
3 FILLED PAUSE PRODUCTION
Filled-pause frequency is shaped by speech rate, age, sentiment, and power status, with gender effects reversing between South and West Slavic parliaments. The study uses negative binomial GEE models and highlights sample imbalance and unresolved explanations as interpretive boundaries.
- Predictors: −36% per SD globally: faster speech yields substantially fewer filled pauses across all four parliaments.The model predicts filled-pause frequency per utterance with a log-duration offset.
- Predictors: −16% per decade globally: older speakers produce fewer filled pauses, significantly so in Czechia and Serbia.The direction is consistent in Croatia and Poland but non-significant there.
- Predictors: +6% per unit globally: more positive speech corresponds to slightly more filled pauses.This association is stable across the full corpus.
- Other political predictors: −21% globally: opposition speakers show a filled-pause discount, significant in Czechia and Croatia but not Poland.The effect is borderline reversed in Serbia, while political orientation has no robust global effect.
- Gender differences: −47% in Croatia and −60% in Serbia: male speakers produce fewer filled pauses than female speakers in South Slavic parliaments.Czech and Polish parliaments show no significant gender difference, reversing the usual conversational-corpus pattern.
- Scope and caveats: Women are a numerical minority among speakers in all four legislatures, widening confidence intervals for the less-represented group.The authors caution that this imbalance matters when interpreting gender-effect size, not only significance.
- Scope and caveats: The data cannot adjudicate whether the gender split reflects parliamentary culture, floor-time and seniority, or broader public-speaking norms.These explanations are presented as non-exclusive and motivate further political-rhetoric research.
4 PRIMARY STRESS IN CROATIAN
The Croatian analysis asks whether speakers’ early-versus-late stress preferences remain consistent across parts of speech. Preferences cohere for verbs, adjectives, and nouns but decouple for adverbs and proper nouns.
- Research question: The analysis addresses whether speaker-level stress preferences cohere across the lexicon or fragment by part of speech.Croatian permits multiple legitimate stress positions for the same multisyllabic word, motivating category-level comparison.
- Data and measurement: Every Croatian multisyllabic word was automatically annotated for primary stress using a speech transformer with 99% word-level accuracy.The analysis uses 11.3 million multisyllabic tokens from 496 speakers and focuses on Croatian because its annotation pipeline was first validated at scale.
- Cross-category coherence: r = 0.88*** for adjectives and r = 0.84*** for nouns: verbs, adjectives, and nouns form a tight central cluster of speaker preferences.Each speaker is represented by their proportion of early-stress usage for each part of speech.
- Category-specific decoupling: r = 0.63–0.70***: proper nouns are only partially aligned with the verbs–adjectives–nouns cluster.Their relationship with the open-class core is weaker than the core categories’ mutual alignment.
- Category-specific decoupling: r = 0.23–0.27***: adverbs lie largely outside the central cluster, while adverbs and proper nouns are nearly independent at r = 0.04, n.s.The MDS projection summarizes these pairwise correlations in two dimensions.
5 FUTURE WORK
Future work extends the corpus, annotations, phonetic analyses, disfluency modelling, stress-variation research, and study of political rhetoric. Key priorities include adding languages and pause types, testing cross-linguistic robustness, and modelling institutional and rhetorical variation.
- Corpus Expansion and Additional Annotation Layers: Slovenian is the immediate next language addition, followed by Bosnian, additional Serbian varieties, Bulgarian, and Ukrainian.These additions offer typological or sociolinguistic leverage, with Slovenian methodologically tractable because of its linguistic proximity to the existing languages.
- Corpus Expansion and Additional Annotation Layers: Silent pauses and finer filled-pause phonetic subclasses would complement the existing hesitation annotations.Silent pauses can be extracted from word-level alignments, while filled pauses could be subclassified by phonetic type.
- Corpus Phonetics: Corpus phonetics can replace small-sample read-speech descriptions with large naturalistic measurements across the four represented languages.Available grapheme- and phoneme-level alignments make analyses of vowel spaces, voice onset time, and consonant-cluster simplification tractable.
- Disfluencies: Future disfluency work should test syntactic positions, surprisal, filled-to-silent pause ratios, and whether filled-pause predictors transfer to silent pauses.The existing annotation and timing layers support analyses of pause clustering, co-occurrence, and inter-word timing.
- Primary Stress Variation: Stress research should replicate Croatian speaker-level coherence patterns in Serbian and examine finer grammatical subcategories.Further work also proposes classifying pitch-accent and stress systems and testing lexical specificity, accommodation, and overcorrection.
- Political Rhetoric: Political-rhetoric research should test whether pitch, rate, and intensity vary with policy domain or government–opposition status beyond sentiment.These questions treat acoustic and disfluency signals as potential rhetorical resources in parliamentary speech.
- Research Scope: Future analyses require sociolinguistic and historical expertise to model institutional cultures and procedural heterogeneity.A key unmodelled boundary is the balance between scripted delivery and extemporaneous response across parliaments, parties, and debate types.
6 CONCLUSION
Using a shared corpus and methodological toolkit, the paper studies sentiment acoustics, hesitation behaviour, and lexical coherence in stress preferences. The findings show cross-linguistic regularities alongside structured differences, supporting comparative research on language, affect, and political behaviour.
- Conclusion: The three studies examine acoustic realisation of sentiment, predictors of hesitation, and coherence of phonological preferences across the lexicon.Together they address cross-modal, behavioural, and structural questions in parliamentary speech.
- Conclusion: Negative sentiment produces a consistent higher, louder, faster acoustic profile across all four parliaments, with an upturn at the most positive scale extreme.The conclusion also reports sharp South–West Slavic divergence along the gender axis of hesitation predictors.
- Conclusion: Croatian stress preferences cohere among verbs, adjectives, and common nouns but fragment for adverbs and proper nouns.This complicates accounts that treat a speaker’s stress system as unitary.
- Conclusion: The results support cross-linguistic designs that treat language-level variation as part of the signal rather than controlling it away.The paper presents its findings as preliminary in a useful sense because each study opens further research questions.