Source-linked AI summary

Understanding State Preferences With Text As Data: Introducing the UN General Debate Corpus

Alexander Baturo, Niheer Dasandi, Slava J. Mikhaylov

arXiv:1707.02774v1cs.CLcs.AIstat.ML

TL;DR

Research on government preferences has relied on limited observable behaviors, while General Debate speeches offer information across issues and over time. The paper introduces the UNGDC and demonstrates text-analytic estimation of state positions, showing how the corpus can complement existing preference measures and support further study of their relationship to UNGA voting. The paper’s scope is bounded by the limited number of issues available in UNGA voting data.

  • Problem

    Existing estimates of state preferences rely on limited observed behaviors and, in UNGA voting, the limited set of issues voted on in a given year.

  • Method

    The paper introduces the UNGDC and uses text analytic methods on General Debate statements to derive estimates of government preferences.

  • Results

    The UNGDC provides estimates that complement existing measures of government preferences and can be compared with UNGA voting behavior across issue areas.

  • Takeaways & Limitations

    The corpus supplies an additional source for studying the nature, formation, and effects of state preferences in world politics.

  • Takeaways & Limitations

    UNGA voting-based preference estimates cover only the limited number of issues voted on in a given year.

Abstract

from arXiv · show

Every year at the United Nations, member states deliver statements during the General Debate discussing major issues in world politics. These speeches provide invaluable information on governments' perspectives and preferences on a wide range of issues, but have largely been overlooked in the study of international politics. This paper introduces a new dataset consisting of over 7,701 English-language country statements from 1970-2016. We demonstrate how the UN General Debate Corpus (UNGDC) can be used to derive country positions on different policy dimensions using text analytic methods. The paper provides applications of these estimates, demonstrating the contribution the UNGDC can make to the study of international politics.

Introduction

The paper addresses the limited use of speeches to measure government preferences in international relations by introducing text analysis of annual General Debate statements and the UNGDC dataset. Because these statements cover many issues and countries over time, they provide additional measures for studying international politics.

  • UN General Debate statements offer largely untapped information on governments’ policy preferences across issues and over time.
  • Existing preference measures rely heavily on military alliances or UNGA votes, which provide limited information outside alliances and cover only the issues voted on in a given year.
  • Text analysis of General Debate statements can broaden measures of government preferences and their effects.
  • The General Debate is suitable for estimating state preferences because all UN member states receive an equal opportunity to speak in a formal, annual setting.
  • The UNGDC contains 7,314 General Debate statements from 1970–2014, prepared for empirical applications using text-as-data methods.
  • The paper demonstrates how the corpus can derive government-preference estimates and support future quantitative research on international politics.

The UN General Debate and world politics

The General Debate gives governments an annual public forum to discuss international issues and express priorities with fewer external constraints than in UNGA voting. Its speeches remain strategic, but their content can reveal national priorities, issue salience, and changing diplomatic positions.

  • The General Debate begins each UNGA regular session annually and provides nearly two hundred member states an opportunity to present their views.
  • Speeches address conflicts, terrorism, development, climate change, non-proliferation, and other shared issues in international politics.
  • Unlike UNGA roll-call votes, General Debate speeches are not institutionally connected to UN decision-making, reducing external constraints on speakers.
  • The lower constraints allow governments to emphasize important issues and provide more information about national priorities than the limited number of annual UNGA votes.
  • Interviews with diplomatic representatives describe General Debate speeches as a place where states can express their views and identify important issues.
  • Despite fewer external constraints, governments use speeches strategically to signal preferences and influence perceptions of their states and others.
  • The changing US rhetoric toward Iran between 2012 and 2013 illustrates the strategic nature of General Debate speeches.
  • Text analysis of General Debate speeches can identify the issues governments emphasize and the important topics emerging in international politics over time.

UNGDC: The UN General Debate Corpus

The UNGDC is a processed corpus of English-language General Debate statements collected from UN sources for 1970–2014. It covers nearly the full membership over time and records substantial variation in speakers and document length.

  • The researchers collected speeches from UN General Debate webpages and UN bibliographic sources, using optical character recognition for earlier image-based documents.
  • Non-English speeches were represented through official English translations supplied by the United Nations.
  • The corpus contains 7,314 country statements delivered from 1970–2014.
  • On average, speeches contain 123 sentences and 945 unique words.
  • Among classified statements, 44.3% were delivered by heads of state or government, 49.3% by senior ministers, and 6.4% by country representatives.

Empirical application: Preferences on single issue dimensions

The paper shows how the UNGDC can make government speeches accessible for text-as-data research and estimate positions on specific policy dimensions. A Wordscore application uses US and Russian statements as references to map differences among states in the 2014 General Debate.

  • The UNGDC provides accessible statements for reading, comparing, and quantitative research on the nature, formation, and effects of state preferences.
  • Text-as-data methods represent documents by word frequencies and use supervised reference texts to define a substantive policy dimension.
  • The application defines a Russia-versus-USA dimension using US and Russian statements as reference texts and estimates positions for the 2014 General Debate.
  • The displayed scores are based on standard preprocessing and LBG rescaling, so predicted values may fall outside the −1 to +1 reference range.
  • The paper does not use these estimates as an explanatory variable in an empirical application because of limited space.
  • The resulting scores demonstrate that text data can estimate differences between UN member states.

Empirical application: Preferences on multiple dimensions

The paper uses correspondence analysis to estimate countries’ relative emphasis across multiple policy dimensions, then illustrates its value in explaining positions on US nonsurrender agreements to the ICC. Three CA dimensions are selected by cross-validation; the third dimension is associated with security and terrorism concerns, although fuller analysis is needed.

  • Method: Correspondence analysis estimates countries’ relative emphasis on multiple issues from unique-word counts, with dimensions interpreted inductively after estimation.The dimensions may represent single issues, multiple issues, or meta-issues rather than predetermined policy scales.
  • Application: The analysis extends an existing study of US nonsurrender agreements by adding multidimensional CA estimates to the model.The application addresses preferences that are unlikely to be reduced to a single issue dimension.
  • Model selection: Leave-one-out cross-validation selects three CA dimensions from specifications allowing up to ten dimensions.The selected dimensions are added to the original model predicting whether countries signed nonsurrender agreements.
  • Results: The CA3 coefficient is statistically significant, and its defining words suggest that stronger security and terrorism concerns were associated with signing the agreement.The authors interpret this pattern as indicating that security concerns alongside normative goals influenced signing decisions.
  • Scope: The authors limit the application’s substantive discussion and state that further analysis is required to fully support the interpretation.The example is illustrative rather than a comprehensive treatment of the issue.

Conclusion

The UNGDC provides a publicly available basis for measuring single and multiple dimensions of government preferences from General Debate statements and applying those estimates to international politics. These estimates complement existing UNGA voting measures and support research on how preferences are expressed, compared, and formed.

  • The UNGDC is a new dataset for understanding and measuring state preferences in world politics.
  • Text analytic methods can uncover both single and multiple dimensions of government preferences from UNGDC statements.
  • UNGDC-derived estimates can be applied to questions about government preferences and their effects.
  • UNGDC estimates complement existing measures based on UNGA voting and can be used to investigate relationships between expressed preferences and voting behavior.
  • Text provides detailed information about countries’ views on particular policy areas, enabling comparisons with international treaties and laws.
  • Such comparisons can examine countries’ influence on international agreements, perceptions of agreements, and adoption of international-law language in General Debate statements.
  • The UNGDC can also help researchers study how state preferences are formed and which domestic groups influence preferences across issues.
Loading 1707.02774v1…