Source-linked AI summary

A Multi-Axis Annotation Scheme for Event Temporal Relations

Qiang Ning, Hao Wu, Dan Roth

arXiv:1804.07828v2cs.CL

TL;DR

Existing TempRel annotation schemes often produce low agreement even among experts, while annotation is also labor intensive. The paper introduces multi-axis modeling and start-point-focused annotation, then uses the better-defined scheme to support crowdsourcing. Expert pilot annotation reached .84 Cohen’s Kappa, compared with conventional values in the 60s.

  • Problem

    Existing TempRel datasets often have expert inter-annotator agreements below or near 60%, indicating a difficult annotation task, while collecting annotations is labor intensive.

  • Method

    The paper models events on separate semantic axes and annotates temporal relations using start-points while excluding end-point comparisons.

  • Results

    .84 Cohen’s Kappa was achieved by expert annotators on a subset of TB-Dense, compared with conventional values in the 60s.

  • Takeaways & Limitations

    The proposed scheme supports crowdsourcing to collect a new dataset and is associated with improved baseline performance on that dataset.

  • Takeaways & Limitations

    The approach ignores event end-point comparisons, leaving event duration as a separate task for further investigation.

Abstract

from arXiv · show

Existing temporal relation (TempRel) annotation schemes often have low inter-annotator agreements (IAA) even between experts, suggesting that the current annotation task needs a better definition. This paper proposes a new multi-axis modeling to better capture the temporal structure of events. In addition, we identify that event end-points are a major source of confusion in annotation, so we also propose to annotate TempRels based on start-points only. A pilot expert annotation using the proposed scheme shows significant improvement in IAA from the conventional 60's to 80's (Cohen's Kappa). This better-defined annotation scheme further enables the use of crowdsourcing to alleviate the labor intensity for each annotator. We hope that this work can foster more interesting studies towards event understanding.

1 Introduction

TempRel annotation has remained difficult and labor intensive, with expert agreements often near 60%. The paper proposes multi-axis modeling, start-point-focused annotation, and crowdsourcing to improve annotation quality and scalability.

  • Expert-annotated TempRel datasets such as TB-Dense, RED, and THYME-TimeML report inter-annotator agreements below or near 60%.
  • Difficult event pairs can support conflicting temporal interpretations, especially when events involve attempts, expectations, or hypothetical situations.
  • The paper introduces multi-axis modeling so annotators compare events only when they belong to the same semantic axis.
  • The proposed scheme focuses TempRel comparisons on start-points because event end-points are especially difficult to annotate.
  • .84 Cohen’s Kappa was achieved by expert annotators in a pilot study on a subset of TB-Dense, versus conventional values in the 60s.
  • The work also facilitates crowdsourcing for collecting a new high-quality TempRel dataset and reports improved baseline performance on it.

2 Temporal Structure of Events

Existing TempRel schemes force annotations for many poorly defined event pairs, producing low agreement. The paper argues for multi-axis modeling that restricts comparisons to semantically comparable events while representing intentions, opinions, and hypotheses separately.

  • Motivation: Dense schemes force annotators to label every event pair in a window, including pairs whose temporal relation is not well-defined.In one example, six pairs are presented even though some pairs are explicitly described as ill-defined.
  • Motivation: Ambiguous durations, hypothetical events, and competing interpretations make difficult pairs produce disagreement even when annotators mark them as vague.An 80% confidence rule does not resolve the problem because annotators can interpret that confidence threshold differently.
  • Multi-Axis Modeling: Multi-axis modeling assigns events to semantic axes and compares only events on the same axis, avoiding forced decisions about often vaguely defined cross-axis pairs.The proposed model includes axes for intentions, opinions, and hypotheses alongside an article’s main axis.
  • Multi-Axis Modeling: The annotation process classifies event anchorability and then labels every pair of anchorable events one axis at a time.The paper treats ruling out cross-axis relations as a strategy for separating well-defined from ill-defined relations, not as a claim that cross-axis relations are unimportant.
  • Orthogonal Axes: Orthogonal axes can connect through intersection events, but projected relations remain valid only under the corresponding opinion or other axis condition.For example, a projected “hardest hit” relation is valid only when the relevant opinion is true; intentions and opinions may also be unfulfilled or false.

3 Interval Splitting

Interval splitting replaces a single interval relation with explicit point comparisons, preserving expressivity while enabling more precise treatment of vague relations. The paper focuses on start-point comparisons because end-point relations are substantially harder for annotators.

  • Existing schemes reduce Allen’s 13 interval relations, whereas interval splitting explicitly compares event time points using before, after, and equal labels.The event intervals are decomposed into four point comparisons, while the relation set remains simple.
  • Interval splitting preserves information in vague cases when annotators agree on a start-point relation but disagree about the full interval relation.This information would otherwise be collapsed into a vague interval label.
  • 4 point comparisons vs. 1 interval comparison is the explicit annotation-cost trade-off of interval splitting.In practice, fewer than four comparisons are often needed, and some comparisons can be inferred.
  • Ambiguity of End-Points: 11% of crowdsourcers passed the qualifying test for end-point relations, while end-end accuracy fell from 67% to 37%.The average response time also increased from 33s for start-start relations to 52s for end-end relations.
  • Ambiguity of End-Points: End-point comparisons are attributed to ambiguity in how durative-event durations are expressed and perceived in natural language.The paper therefore ignores end-point comparisons despite treating event duration as important.

4 Annotation Scheme Design

The annotation scheme first restricts comparisons through anchorability and dense labeling, then uses quality controls and paired questions to reduce inconsistent or vague decisions.

  • The two-step scheme marks verb event candidates as Anchorable or not, then labels TempRels only between Anchorable events.Non-verb event candidates are removed during preprocessing.
  • Crowdsourcing uses expert-annotated gold examples for a 70% qualifying test on 10 questions and an undisclosed surviving test during annotation.These protocols are implemented through CrowdFlower’s quality-control feature.
  • Vague Relations: For start-point relations, the label set is before, after, equal, and vague, with two possibility questions mapped to these four labels.The mapping is vague for two yes answers, equal for two no answers, and directional otherwise.
  • Vague Relations: Asking both possibility questions prompts annotators to consider all possibilities, reducing the chance of overlooking a relation.

5 Corpus Statistics and Quality

The pilot and crowdsourced annotations show substantially improved agreement under MATRES, while comparisons with TB-Dense support consistency between the datasets despite their different annotation targets.

  • Expert pilot: κ=.85 was achieved for anchorability, followed by κ=.90 for Q1 and κ=.87 for Q2 in the expert pilot on the main axis.Only events labeled Anchorable by both experts were retained for relation annotation.
  • Expert pilot: Vague relations had lower agreement than other labels, with κ=.75 and F1=.81, confirming that temporal vagueness remained more difficult.
  • Crowdsourcing: 28% of events were labeled Non-Anchorable by crowdsourcers, with gold accuracy .86 and WAWA .79.
  • Crowdsourcing: Crowdsourcers showed strong agreement with the gold annotations and among themselves, indicating that the relation task was understandable to non-experts.The two measures were gold-set accuracy and Worker Agreement with Aggregate (WAWA).
  • Additional axes: For INTENTION and OPINION axes, crowdsourcers achieved gold accuracy .82 and WAWA .89, while two experts reached .86 F1 on the second step.These axes contained 16% of events and were usually short.
  • Corpus construction: Each MATRES judgment cost $0.01, and the 36-document dataset cost about $400 overall.
  • Corpus comparison: MATRES contains 0.8K main-axis events yielding 1.6K EE relations and 0.2K orthogonal-axis events yielding 0.2K EE relations, from 1.1K TB-Dense verb events.The comparison with TB-Dense uses 1.8K EE relations in common after mapping interval annotations to start-point relations.
  • Corpus comparison: TB-Dense and MATRES agree strongly on before, after, and vague labels, while MATRES often resolves TB-Dense vague relations into before, after, or equal.The reported agreements are b=.89, a=.71, and v=.62; the difference is attributed to interval versus start-point annotation.

6 Baseline System

The baseline classifies event-pair relations with an averaged perceptron using lexical, syntactic, distance, modal, and connective features. Retraining on MATRES confirms that the annotation scheme produces a better-defined learning task than the original TB-Dense annotations.

  • System: The baseline assumes that events and axes are given and predicts one of four relation types for each event pair.
  • Features: Features include POS tags, event distance, intervening modals, temporal connectives, and shared verb synonyms.The features use each event and neighboring words together with sentence and token distances.
  • Training: The system uses an averaged perceptron and the TB-Dense split of 22 training, 5 development, and 9 test documents.Parameters were tuned on the training set and development set before retraining on their union.
  • Results: The same system was retrained and tested on original TB-Dense annotations as an “Original” comparison, confirming significant improvement with the proposed annotation scheme.The paper attributes the improvement to the scheme and resulting dataset rather than claiming the baseline algorithm itself is superior.

7 Conclusion

The paper proposes multi-axis, start-point-focused TempRel annotation to address low agreement and endpoint confusion. Expert pilots and crowdsourced evaluations support the scheme and its MATRES dataset.

  • Contributions: The scheme simplifies TempRel annotation by focusing on one semantic time axis at a time and comparing only events on the same axis.
  • Contributions: Because event endpoints create major confusion, the paper proposes annotating start-points and leaving endpoint issues for future duration-annotation investigation.
  • Evidence: Expert pilots show significant IAA improvement over literature values, indicating a better-defined annotation task.
  • Dataset: The scheme enables crowdsourcing MATRES at lower time cost, with quality supported by gold-set agreement, WAWA, and consistency with TB-Dense.
  • Outlook: The authors hope these findings provide a starting point for understanding more sophisticated semantic phenomena in event understanding.

A Examples of Table 1

The examples illustrate categories in the proposed temporal structure, including orthogonal and parallel axes and events excluded from any axis. Figure 4 lists thirteen interval relations.

  • Orthogonal axis: INTENTION is defined as the actual intent of verbs, allowing verbs outside TimeML’s I-Action and I-State categories to create an orthogonal intention axis.The example includes “allocated” as an intention event.
  • Examples: Example 8 places leave, build, and win on INTENTION or OPINION orthogonal axes, while elected, cut, contains, and hunt illustrate parallel-axis categories.
  • Examples: NEGATION events such as helping and want are shown as not belonging to any axis.
  • Interval relations: Figure 4 enumerates thirteen possible relations between two event intervals, including before, after, equal, starts, ends, includes, and overlap variants.

B Anchorable vs. Actual

The paper compares Anchorable with Actual labeling on 166 verb events from RED and finds that the categories differ substantially, especially for several event types.

  • Two experts achieved Cohen’s Kappa of .88 for anchorability annotation on 166 verb events from five randomly selected RED documents.This matched their Cohen’s Kappa on MATRES.
  • Anchorable and Actual labeling differed statistically on the RED subset, with McNemar’s p ≪0.001.
  • The comparison found no Anchorable events labeled Non-Actual, but 25 events labeled Actual by RED were Non-Anchorable here.The paper attributes the absence of Anchorable/Non-Actual cases to their lower practical frequency.
  • Among the 25 discrepant cases, 11 were INTENTION, 4 OPINION, 6 STATIC, and 4 NEGATION.
  • Examples of Non-Anchorable cases included compensation as INTENTION, sponsoring terrorism as OPINION, based as STATIC, and subsidize under NEGATION.

C Annotation Interface

The annotation interface separates anchorability decisions from relation decisions and splits the two relation questions into separate tasks to avoid answer-ordering bias.

  • In the first step, crowdsourcers judge one event at a time in full context using a binary Yes/No anchorability decision.
  • The second step asks two questions for each event pair to determine the actual TempRel.
  • Showing both questions simultaneously wrongly suggests that annotators should choose one “yes” and one “no,” creating strong correlation between answers.
  • The final design launches separate tasks for Q1 and Q2, preventing one annotator from seeing both questions simultaneously.This is intended to make annotators consider the temporal relation rather than automatically selecting opposite answers.
  • The owner can inspect crowdsourcers’ answer distributions, such as 86% and 14%, although crowdsourcers cannot see those distributions.
Loading 1804.07828v2…