Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX

Paper · arXiv 2609.18011 · Published September 16, 2026
NLP and Linguistics

In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evidence about grounding across two such tasks. Working from discrete behavioral annotations, we map HCRC MapTask (Anderson et al., 1991) and MUNDEX (Türk et al., 2023) into a shared partner/task/away vocabulary and compute gaze features around task-relevant dialogue units. In both corpora, aligned reference interpretations (MapTask) and UND (understood) judgments (MUNDEX) are associated with more task-directed gaze and with less partner-directed gaze, lower gaze entropy, and fewer gaze transitions. The associations are clearest for the participant leading the task: in giver-produced references, and in explainer judgments, which also co-vary with the explainee’s gaze. In same-speaker MapTask reference chains, the speaker’s gaze entropy is lower at the mention where a previously nonaligned referent becomes aligned. The best gaze feature groups improve modestly over controls under grouped cross-validation: temporal features in MapTask and raw proportions in MUNDEX.

Introduction. In collaborative tasks where participants hold different private information, mutual understanding cannot be assumed from shared context alone. It must be built and tracked through interaction (Clark and Wilkes-Gibbs, 1986; Clark and Brennan, 1991). Gaze is an observable cue to this process: participants look at task materials, at each other, or away while giving instructions, checking understanding, and coordinating their perspectives. Some corpora annotate gaze from video as discrete categories of where participants look, rather than as eye-tracking coordinates. These annotations can be used to study the relationship between gaze and grounding, but they are often corpus-specific, making it difficult to compare across tasks. We study two settings where information asymmetry forces participants to continuously coordinate understanding. In HCRC MapTask (Anderson et al., 1991), a giver and a follower navigate with maps that differ in their landmarks; perspectivist grounding labels record each participant’s interpretation separately (Li et al., 2026a).

Discussion / Conclusion. Shared categories, task-specific meanings The direction of these associations is the same in both corpora, echoing map-task observations that partner-directed gaze increases around communicative difficulty (Boyle et al., 1994; Nakano et al., 2003; Murat and Vogel, 2026). The two labels measure different constructs: MapTask records referential alignment, whereas MUNDEX pools explainees’ self-reports and explainers’ judgments, so the convergence spans related but distinct grounding measures. Which features carry predictive signal differs: temporal dynamics score highest in MapTask and raw proportions in MUNDEX. The shared categories also name gaze targets rather than functions. In MapTask, a partner glance may check a landmark reference; in MUNDEX, gaze averted from the partner has also been linked to topic changes (Lazarov and Grimminger, 2026), so it Gaze and interactional role Significant associations concentrate in giver-produced references and explainer judgments, whereas follower-produced references show near-zero effects.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How does dialogue structure affect linguistic grounding and shared meaning? Should agents decouple planning from perception grounding for better performance? Can multi-agent systems avoid converging on false agreement without deliberation? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? Do language models reason like humans or mimic surface patterns? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Why is dynamic grounding necessary for achieving true mutual understanding in dialogue? What design and behavioral factors drive false consciousness attribution to AI? Can AI systems distinguish genuine empathy from simulated emotion? What happens to knowledge when intelligence becomes tokenized like a commodity? Can language models build genuine grounding through interaction? Does preference optimization systematically degrade conversational grounding in language models?