Source-linked AI summary

Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX

Nan Li, Albert Gatt, Massimo Poesio

arXiv:2609.18011v1cs.CLcs.HC

TL;DR

Collaborators with asymmetric information must establish mutual understanding through interaction, and the paper asks whether gaze provides evidence of that grounding. It maps MapTask and MUNDEX into a shared partner/task/away vocabulary and tests gaze around grounding-relevant dialogue units. Aligned references and UND judgments are associated with more task gaze and less partner gaze, especially for task-leading participants, but effects are small and sensitive to inference unit.

  • Problem

    In collaborative tasks with asymmetric information, mutual understanding cannot be assumed from shared context, while corpus-specific discrete gaze annotations make cross-task comparison difficult.

  • Method

    The study maps MapTask and MUNDEX discrete gaze annotations into a shared partner/task/away vocabulary and analyzes gaze features around grounding-relevant windows.

  • Results

    Aligned MapTask references and MUNDEX UND judgments are associated with more task-directed gaze, less partner-directed gaze, lower entropy, and fewer transitions, clearest for givers and explainers.

  • Takeaways & Limitations

    Gaze provides directionally consistent evidence about grounding and should be modeled alongside linguistic, dialogue, and task-context features.

  • Takeaways & Limitations

    Small samples, recurring participants, modest effects, and sensitivity to inference units limit conclusions about dynamic grounding.

Abstract

from arXiv · show

In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson et al., 1991) and MUNDEX (Türk et al., 2023) into a shared partner/task/away vocabulary and compute gaze features around task-relevant dialogue units. In both corpora, aligned reference interpretations (MapTask) and UND (understood) judgments (MUNDEX) are associated with more task-directed gaze and with less partner-directed gaze, lower gaze entropy, and fewer gaze transitions. The associations are clearest for the participant leading the task: in giver-produced references, and in explainer judgments, which also co-vary with the explainee's gaze. In same-speaker MapTask reference chains, the speaker's gaze entropy is lower at the mention where a previously non-aligned referent becomes aligned. The best gaze feature groups improve modestly over controls under grouped cross-validation: temporal features in MapTask and raw proportions in MUNDEX. Because effects are small and several weaken when recurring participants rather than dialogues are the unit of inference, we treat gaze as one contributing cue to grounding, to be interpreted alongside task and dialogue context.

1 Introduction

The paper asks whether gaze provides evidence of grounding when collaborators hold asymmetric information. It compares MapTask and MUNDEX using a shared gaze representation to study grounding-related patterns across tasks.

  • 1 Introduction: The study compares grounding-related gaze patterns across MapTask and MUNDEX, two tasks requiring continuous coordination under information asymmetry.MapTask records referential alignment, while MUNDEX records explainees’ self-reports and explainers’ judgments of understanding.
  • 1 Introduction: A shared partner/task/away representation enables comparison of corpus-specific gaze annotations across tasks and other video-coded corpora.The representation maps discrete behavioral categories rather than eye-tracking coordinates.
  • 1 Introduction: The paper tests whether gaze patterns converge directionally across distinct grounding measures and whether associations differ by interactional role.It also examines gaze entropy within same-speaker MapTask reference chains and characterizes association strength with grouped prediction and role-stratified tests.

2 Related Work

Prior work shows that gaze can signal understanding, communicative difficulty, and referential information in collaborative tasks. The paper extends these observations through a shared cross-corpus gaze vocabulary spanning MapTask and MUNDEX.

  • 2 Related Work: Gaze has been linked to builders’ displays of understanding, directors’ utterance adjustments, and target identification before linguistic disambiguation.These findings motivate treating gaze as evidence available during collaborative grounding and reference resolution.
  • 2 Related Work: A MapTask example shows a pending follower reference accompanied by repeated partner glances, followed by alignment after the giver’s re-mention.Both participants remain predominantly task-directed in the illustrated excerpt.
  • 2 Related Work: MapTask studies report more partner-directed gaze around differing landmarks and dialogue-act-specific differences between partner and map attention.Sustained speaker gaze after an assertion was usually followed by elaboration, whereas continued map attention more often preceded the next instruction.
  • 2 Related Work: MUNDEX studies relate understanding to speaker information value, syntactic complexity, listener gaze entropy, and manually coded partner, table, or away gaze.The present work builds on these within-corpus analyses rather than assuming identical meanings for the two tasks’ grounding labels.
  • 2 Related Work: The paper maps corpus-specific gaze annotations into one partner/task/away vocabulary so gaze associations can be compared across grounding measures.MapTask provides reference-level labels that support comparisons across repeated mentions of the same landmark.

3 Data and Representation

The study combines gaze-annotated MapTask and MUNDEX data with a shared discrete partner/task/away representation. Grounding labels and gaze windows are matched within each corpus and filtered for sufficient coverage.

  • 3 Data and Representation: MapTask alignment requires speaker and addressee interpretations to resolve to the same landmark, with separate eye-contact and no-eye-contact conditions.The up→partner mapping is literally partner-directed only in the eye-contact condition.
  • 3 Data and Representation: After coverage filtering, MapTask contributes 5,144 reference-expression windows: 3,807 aligned, 1,261 pending, and 76 misunderstood.The sample comes from 46 dialogues with gaze annotations for both participants; 17 windows were excluded for insufficient coverage.
  • 3 Data and Representation: MUNDEX records explainers teaching a board game and explainees retrospectively reporting understanding, while explainers judge explainee understanding on four levels.The levels are understood, partially understood, not understood, and misunderstood.
  • 3 Data and Representation: After filtering, MUNDEX provides 807 gaze windows from 956 valid annotations, including 360 UND, 199 PART_UND, 151 NON_UND, and 97 MISUND windows.The retained windows comprise 458 explainer and 349 explainee observations.
  • 3 Data and Representation: The shared representation maps MapTask up/down/off and MUNDEX EX/EE/TABLE/AWAY into partner, task, and away categories.Non-positive-duration gaze events are discarded and temporal overlaps are resolved within each participant’s stream.

4 Experiments

The experiments test associations between gaze features and grounding states, analyze gaze changes across repeated references, and evaluate whether gaze supports grouped prediction. Features range from target proportions to temporal and coordination measures.

  • 4 Experiments: Gaze features are computed per window from raw proportions, coverage, transitions, entropy, temporal dynamics, transition bigrams, coordination, and derived ratios.The shared vocabulary supports both participant-level and joint gaze features.
  • 4 Experiments: Association tests compare aligned with non-aligned MapTask windows and UND with non-UND MUNDEX windows using rank-biserial correlations, robust GEE, and BH-adjusted q values.Analyses are stratified by interactional role and, in MapTask, eye-contact condition.
  • 4 Experiments: The process analysis tracks gaze across repeated mentions of the same landmark to assess within-speaker changes around alignment.The main results are reported in Table 1 and Figure 2, with full tests in the appendices.
  • 4 Experiments: Prediction binarizes MapTask alignment and MUNDEX UND labels, then evaluates standardized balanced logistic regression with grouped cross-validation.MapTask uses 10-fold dialogue grouping and MUNDEX uses 5-fold grouping.
  • 4 Experiments: The feature groups test whether gaze contains recoverable signal beyond controls under corpus-specific binary grounding labels.Minority classes motivate binarization after intersecting labels with gaze-coverage requirements.

5 Results

Across both corpora, grounding-related gaze associations point toward more task-directed and less partner-directed gaze, with lower entropy and fewer transitions, but effects are small and scope-dependent.

  • Task gaze increased while partner gaze, entropy, and transitions generally decreased for aligned MapTask references and MUNDEX UND judgments.The largest pooled associations were |r|=.058 in MapTask and |r|=.181 in MUNDEX.
  • The strongest associations were role-specific: MapTask effects were clearest for giver-produced references, while MUNDEX explainer judgments had larger effects than explainee self-reports.Giver-produced references had six significant features; MUNDEX effects were |r|=.206 for EX judgments versus .151 for EE self-reports.
  • In MapTask, all features were larger in the eye-contact condition, whereas no feature was significant without eye contact, although the formal interaction was not significant.The largest |r| was .078 with eye contact versus .025 without it.
  • Aligned rates rose from .30 at first mentions to .59 at second and .85 at fourth-or-later mentions, alongside lower partner gaze, entropy, and transitions at second mentions.
  • At the first subsequent aligned mention in same-speaker chains, speaker entropy decreased (dz=−.20, q=.044), but the effect disappeared with dialogue-level aggregation (q=.20).Partner gaze, task gaze, and transitions shifted similarly but did not survive correction; only second-mention task gaze survived the same-position comparison (q=.01).
  • Grouped prediction favored structured+temporal features in MapTask (macro-F1 .532) and raw proportions in MUNDEX (.564), improving over controls by .060 and .020.These groups ranked highest in 27 of 30 reshuffled grouped partitions per corpus, but gains remained modest and partition-dependent.

6 Discussion

Across MapTask and MUNDEX, gaze patterns show directionally consistent associations with grounding, although their interpretation depends on interactional role and task context.

  • Aligned references and UND judgments are associated with more task gaze, less partner gaze, lower gaze entropy, and fewer transitions across both corpora.The shared direction spans related but distinct grounding measures rather than identical labels.
  • The strongest predictive feature groups differ by corpus: temporal dynamics rank highest in MapTask, whereas raw proportions rank highest in MUNDEX.Shared gaze categories identify targets, not fixed functions; partner-avoiding gaze in MUNDEX may also organize explanations around topic changes.
  • Task-general and task-specific gaze behavior remain difficult to separate without finer-grained referent and dialogue-action comparisons.In MUNDEX, gaze aversion may reflect explanation organization as well as understanding.
  • Associations concentrate in giver-produced references and explainer judgments, while follower-produced references show near-zero effects.Explainee gaze proportions and dynamics also co-vary with explainers’ judgments, but role interactions do not survive correction.
  • Discrete gaze categories can support cross-corpus modeling alongside lexical content, dialogue acts, and task state.The representation is designed to transfer to other corpora with video-coded gaze annotations.

7 Conclusion

The cross-corpus analysis finds gaze to be a directionally consistent but limited cue to grounding in asymmetric collaborative tasks.

  • Across MapTask and MUNDEX, aligned references and UND judgments coincide with more task gaze, less partner gaze, lower entropy, and fewer transitions.The pattern is clearest for giver-produced references and explainer judgments.

8 Limitations

The findings are constrained by coarse gaze labels, non-equivalent perspectives across tasks, recurring participants, and limited data scope.

  • Representation and labels: Three gaze categories cannot identify the specific landmark or object being viewed, and shared category names do not guarantee equivalent functions across tasks.MapTask’s up-to-partner mapping is literal only when eye contact is possible.
  • Representation and labels: MUNDEX combines explainees’ retrospective self-reports with explainers’ judgments, and linked events can contain conflicting binary labels.Retaining one row per pair preserves the positive association between explainer task-gaze proportion and UND.
  • Evidence and scope: Recurring participants and few groups weaken inference: no MapTask feature survives correction under grouped clustering, while only selected MUNDEX associations remain significant.The modest effects, chain-result sensitivity, and cross-validation sensitivity limit claims about dynamic grounding and predictive generalization.
  • Evidence and scope: The dataset covers 46 MapTask dialogues and 26 MUNDEX interactions with nine explainers, with substantial participant recurrence.Coverage filtering also excludes moments unevenly, and broader task structures and multimodal predictors remain untested.

Ethics Statement

The study analyzes publicly released MapTask and MUNDEX corpora and their annotations, which do not identify individual participants.

  • The study uses only the publicly released HCRC MapTask and MUNDEX corpora, whose annotations do not identify individual participants.
  • MapTask gaze annotations distinguish up, down, and off categories, while its dialogues include eye-contact and no-eye-contact conditions.
  • The analysis resolves gaze events, word timings, and landmark references across annotation layers, yielding 5,162 timed references across 46 dialogues.
  • All 5,161 MapTask grounding annotations matched timed references, leaving 5,144 windows after coverage filtering.
  • Multiple-landmark expressions can duplicate gaze windows, and conflicting labels are handled by retaining a non-aligned observation or excluding conflicting windows.
  • MUNDEX records explainer teaching interactions and retrospective judgments of explainee understanding using UND-related tiers.

A.3 Understanding perspectives and linked annotations

The analysis distinguishes explainer judgments from explainee self-reports and links them to gaze windows while preserving perspective-specific annotations. Prediction results favor raw gaze for explainer judgments and show modest gains over controls overall.

  • Understanding perspectives: EX judgments and EE self-reports remain separate rows, with UND positive and PART_UND, NON_UND, and MISUND negative.This preserves the distinction between the explainer’s judgment and the explainee’s own report.
  • Linked annotations: MUNDEX links EX and EE annotations to UND_MATCH spans only when exactly one valid candidate per perspective is found.Of 164 spans, 159 yield clean links; 135 retain both rows after coverage filtering, producing 271 of 807 rows.
  • Prediction: .564 macro-F1 for MUNDEX raw proportions exceeds the .356 controls-only score, while the role indicator alone reaches .544.Perspective-specific scores are .549 versus .337 for EX judgments and .520 versus .380 for EE self-reports; these are point estimates from overlapping subsets.
  • Prediction: Raw gaze is the highest-scoring feature group for EX judgments, whereas structured+ratios reaches .556 for EE self-reports.The same five-fold held-out-explainer procedure is used for both perspectives.
  • Windows and features: Windows center on understanding annotations with ±2 seconds of context, retain only windows with at least 30% gaze coverage for both participants, and remove non-positive-duration events.Overlapping same-participant events count once for structured features; changing overlap assignment alters no macro-F1 score by .001 or more.
  • Feature construction: The feature inventory combines partner/task/away proportions and mutual gaze with coverage, transitions, entropy, temporal dynamics, bigrams, ratios, and joint coordination measures.Proportions use observed duration over window duration, while entropy uses the distribution of observed gaze time.

D Role- and Condition-Stratified Analyses

Role and inference-unit analyses narrow the gaze–grounding associations: effects are clearest for task leaders and weaken when recurring participants define clusters. Within same-speaker reference chains, entropy decreases at resolution, but this result is sensitive to the inference unit.

  • Role-stratified analyses: Giver-produced MapTask references show small significant effects, whereas follower-produced references show near-zero effects.Effect sizes reach |r| up to .086 for givers and remain at or below .043 for followers.
  • Condition-stratified analyses: All main MapTask associations are significant with eye contact but disappear without it, although the condition interaction is not significant.Mean speaker partner gaze is .04 without eye contact versus .25 with eye contact.
  • Cluster-robust inference: The clustered analyses use separate univariate logistic GEEs because gaze features are strongly collinear.Label proportions, entropy, transitions, and mutual-gaze measures overlap substantially, complicating joint coefficient interpretation.
  • Cluster-robust inference: Participant-level clustering removes BH-significant MapTask features and weakens several MUNDEX associations, showing sensitivity to recurring-participant dependence.Dialogue-cluster corrections retain some effects, but no MapTask feature survives participant-group clustering.
  • Reference-chain analyses: Across 636 landmark chains, same-speaker resolution produces lower speaker entropy, but the result depends on the inference unit.The primary analysis uses 189 same-speaker pairs and finds dz=−.20; dialogue averaging and six-group testing are not significant.

G Prediction Setup and Full Results

Grouped prediction tests whether gaze features generalize beyond controls while role-stratified analyses characterize where associations are concentrated. Gains are modest and vary by task and feature representation.

  • Prediction setup: Grouped logistic regression uses standardized gaze features with balanced class weights and held-out dialogue or explainer folds.MapTask uses 10 dialogue-grouped folds; MUNDEX uses 5 explainer-grouped folds.
  • Role-stratified analyses: MapTask role-stratified tests show significant effects for giver-produced references but not follower-produced references.Giver effects reach |r| up to .086, whereas follower effects remain at or below .043.
  • Full prediction results: Structured+temporal features improve MapTask controls by .038 on average, while raw proportions improve MUNDEX controls by .017.The best feature group ranks first in 27 of 30 reshuffled grouped partitions for each task.
  • Full prediction results: Partition reshuffling shows that MapTask control performance ranges from .460 to .511, while MUNDEX raw-proportion gains are not positive in one partition.This indicates modest but partition-dependent predictive improvements.
  • Role-stratified analyses: MUNDEX role-stratified tests distinguish explainee self-reports from explainer judgments of explainee understanding.The table separates EE and EX strata for interpreting role-specific associations.
Loading 2609.18011v1…