Source-linked AI summary
Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation
Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce
TL;DR
The paper asks whether VLMs can use prior conversational context to interpret ambiguous referring expressions in iterated reference games. It evaluates humans and VLMs under varied context conditions and finds that models use relevant context but struggle to learn from mistakes, remaining substantially worse than humans. The study therefore identifies limitations in the context tracking and adaptation needed for efficient linguistic collaboration.
Problem
The study addresses whether VLMs can perform context-sensitive pragmatic interpretation of referring expressions across multi-turn interactions, a capability important for conversational communication.
Method
Humans and VLMs interpreted descriptions from iterated reference games while the amount, order, relevance, and feedback of prior context were varied.
Results
Models performed ad hoc pragmatic reasoning and used relevant context, but were substantially worse than humans and generally failed to learn from past mistakes.
Takeaways & Limitations
The findings indicate that current VLMs remain limited in bootstrapping their own learning and in the flexible adaptation required for linguistic collaboration.
Takeaways & Limitations
The study used a limited set of abstract stimuli and English responses, and iterated reference games are not a comprehensive measure of pragmatic ability.
Abstract
from arXiv · showhide
Flexible adaptation to context and shared pragmatic intuitions contribute to smooth human conversation. Iterated reference games---in which players repeatedly pick out novel referents using language---present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision--language models on their ability to identify the intended meaning of descriptions produced in iterated reference games, varying the provided context in terms of amount, order, and relevance. While humans performed well consistently, the models we evaluated could make use of prior context to interpret humans' referring expressions, but they struggled to build up the relevant context to interpret those expressions effectively. Our results suggest that the models we evaluated lack core skills needed for efficient linguistic collaboration.
1 Introduction
Multi-turn conversation requires agents to retain relevant prior context and use it to interpret new messages. This study uses iterated reference games to test whether VLMs can pragmatically resolve ambiguous referring expressions across varying contexts, compared with humans.
- Conversational agents must interpret each message using relevant preceding context, alongside natural-language understanding, world knowledge, and instruction following.
- Iterated reference games require a describer to identify a referent through language so a matcher can select it from multiple options.
- Repeated descriptions of the same referent can produce conventionalized expressions and partner-specific meanings that humans adapt to dynamically.
- The study tests whether state-of-the-art VLMs use varying amounts and types of prior context to resolve referential ambiguity, and compares their context sensitivity with humans'.
2 Related work
Pragmatics involves context-sensitive language use across visual, conversational, cultural, and interlocutor-related information. Prior work shows that VLMs face challenges with visual grounding, multi-turn context, and human-like adaptation in reference games.
- VLMs can underuse visual information, with reported weaknesses in low-frequency image classification, shape overlap identification, counting, and spatial reasoning.
- Multi-turn benchmarks report that contemporary models may fail to use earlier relevant information and may produce incoherent responses across conversations.
- Pragmatics encompasses sensitivity to cultural knowledge, grounded visual context, interlocutors' knowledge states, and recent conversational context.
- Iterated reference games combine difficult visual contexts with online adaptation, while prior work finds some success comprehending efficient human descriptions of naturalistic images.
- Human–AI reference-game dyads show lower accuracy than AI–AI and human–human dyads, and models' interaction patterns differ from human patterns.
- Training and explicit prompting can improve performance and make descriptions more superficially human-like, but models have not shown human-like partner-sensitive convention formation.
3 Methods
The study evaluates five instruction-tuned open-weights VLMs and human matchers on tangram reference-game transcripts. Models receive differently ordered, sourced, and informative prior trials, with feedback variants separating retrieval from learning.
- The dataset contains ten games with grids of 12 tangram images, six rounds of 12 trials, and 2–6 players per game.
- Naïve human matchers selected described images from transcripts and received correctness feedback without being told the answer after mistakes.
- Models received 12 labeled tangram options, preceding trials formatted as chat history, and a test description; accuracy was the softmax-normalized probability assigned to the correct target.
- Eight context conditions varied the amount, order, relevance, and fidelity of in-context trials relative to human participants' original games.
- Feedback conditions varied whether models saw their own or human guesses and whether each guess was correct, distinguishing retrieval from learning.
4 Results
Models performed above chance but remained less accurate and less human-like than people, with performance improving mainly when feedback preserved relevant game-specific information. Their behavior suggests retrieval of correct prior answers more than learning from errors or building reusable conventions.
- 4.1 Models struggle to reason correctly in iterated reference games: Models were above chance across cumulative conditions but improved only slightly over rounds, while naïve humans started higher and improved more.Performance was generally poorer in shuffled and random conditions than in yoked conditions for both models and naïve humans.
- 4.2 Providing additional contextual information improves model performance: Human-yoked limited feedback produced substantial improvement, and full feedback brought models close to 100% accuracy in yoked, backward, and shuffled conditions.Under human-yoked limited feedback, models reached an asymptote close to human levels in these conditions.
- 4.3 Only relevant contextual information improves model performance: In non-cumulative conditions, models reached only 0.3–0.6 accuracy when context came from different games, and performance was also low when the test tangram was ablated.The boost therefore depended on relevant in-context trials from the same original game and experience with the test tangram itself.
- 4.4 Models fail to learn from past mistakes: After a correct previous answer, humans and models were likely to answer correctly again; after an incorrect answer, models performed at or below chance while naïve humans corrected about half their mistakes.Follow-up analyses indicated that models perseverated on the same wrong answer rather than learning over repetitions.
- 4.5 VLMs make limited use of visual information: Removing images generally hurt first-repetition performance, but by the third repetition models were usually no different with or without images; performance remained slightly worse without images when relevant context was limited.Some models, especially Llama 3.2, were better without images in many conditions.
- 4.6 Models and humans diverge in their responses: Models showed weak trial-wise similarity to humans, and their response to metaphorical descriptions was much weaker than humans’ response.The estimated effect of metaphorical spans was b = 1.08 for humans versus b = 0.22 for models; holistic-span effects were nonsignificant.
5 Discussion
The models performed ad hoc pragmatic reasoning across multiple turns but remained substantially worse than humans. Their context sensitivity reflected difficulty learning from past mistakes, while performance depended on the amount and relevance of contextual information.
- State-of-the-art open-weights VLMs performed ad hoc pragmatic reasoning across multiple turns but remained substantially worse than humans.
- Models improved more when correct answers were present in context, while learning was limited when correct answers were unavailable.
- The models’ context-use pattern appeared to result from failure to learn from past mistakes, unlike humans.
- Models and humans showed different performance patterns, including weak trial-wise calibration and stronger human accuracy for more metaphorical spans.
- Reference games with abstract referents remain difficult for machine learning models in the few-shot setting, and models remain limited in bootstrapping their own learning.
Limitations
The study’s comparisons were constrained by differences between human and model response procedures and by a limited stimulus and language scope. The reference-game paradigm also does not comprehensively measure pragmatic reasoning.
- Human and model response procedures differed because non-A–L model continuations were ignored, while humans were explicitly restricted to selecting A–L.
- The study tested a limited set of abstract stimuli and focused on English responses, leaving generalisability to naturalistic stimuli and other languages for future work.
- Iterated reference games are not a pure measure of pragmatics or a comprehensive evaluation of pragmatic ability, so complementary tasks are needed.
A.1 Human experiments
Human participants and language models completed reference-game trials using labeled tangram images, conversational transcripts, and feedback. The appendix describes the human instructions, model prompting, inference procedures, and reliability analysis.
- Human participants read transcripts of prior reference-game conversations and selected which of 12 consistently presented images was described.
- Human trials revealed descriptions word-by-word, followed by an image-selection response and correctness feedback.
- Models received the 12 tangram shapes labeled A–L and conversational trials formatted as user text, assistant letter guesses, and correctness feedback.
- Open-weights models used vLLM with bfloat16 precision on two NVIDIA A40 GPUs, while frontier models were queried through their corresponding APIs.
- Model completions were parsed for uppercase letter choices, and Olmo 3 7B Instruct was used to annotate literal/metaphorical and part-based/holistic spans.
- Inter-rater reliability was ICC(A, 1) = .954 [.937, .967] for metaphorical spans and ICC(A, 1) = .824 [.764, .870] for holistic spans.
B.1 Performance varies by original trial number
Performance changed with the original repetition number and varied substantially across items and conditions. Correlation analyses showed systematic clustering among conditions, with relevance organizing the models’ confusion patterns.
- B.1 Performance varies by original trial number: Across shuffled, backward, random, and no-context conditions, performance generally worsened over original rounds as human referring expressions became more idiosyncratic.
- B.1 Performance varies by original trial number: Humans and some models performed best in the second repetition, suggesting practice effects before later conventionalisation made expressions harder to resolve.
- B.2 Item-wise variation: Naïve humans and models showed large item-wise accuracy variation, sometimes exceeding variation due to repetition number, matcher, or condition.
- B.4 Conditions pattern with each other systematically: Under full feedback, model confusion matrices correlated above .85 across most conditions, while ablated and no-context conditions were more distinct.
- B.4 Conditions pattern with each other systematically: Yoked, shuffled, and backward conditions clustered together, as did other-within, other-across, and random conditions, supporting relevance as an organising factor.
B.5 Lemmas are related to model–human discrepancies
The analysis links model–human accuracy discrepancies to the kinds of lemmas used in referring expressions, with literal spatial and geometric terms differing from more metaphorical language. Filtering lemmas by recurrence ensured these discrepancies were not driven by a few idiosyncratic games.
- Discrepancy analysis: Model–human discrepancies were estimated as model accuracy minus human accuracy for each unique trial, averaged across models and runs.The analysis then examined lemmatised descriptions to relate discrepancies to specific words.
- Discrepancy analysis: Lemmas were retained only when they occurred in at least 4 games and at least 20 trials.This filtering reduced the influence of rare, idiosyncratic games.
- Lemma patterns: Geometric and spatial lemmas had positive or zero-centered discrepancy values, whereas more metaphorical words generally required interpretations that produced different model–human patterns.Examples of literal terms include “triangle,” “diamond,” “square,” “left,” “side,” and “top.”
- Lemma patterns: The metaphorical terms “guy” and “man” also had positive discrepancy values, possibly because they were frequent and combined with literal descriptors.The authors note that this pattern converges with their span-based analysis.
B.6 Given additional practice, models still fail to improve
Practice conditions held descriptions constant so models could learn from repeated feedback, but performance remained poor and models increasingly repeated prior answers. Invalid responses were relatively uncommon and were excluded from accuracy and response-probability calculations.
- Practice design: Practice conditions repeated descriptions from a single original round, keeping descriptions fixed while models received multiple chances to learn.The study drew conditions from both the first and last rounds under interactive limited feedback.
- Practice results: Models’ practice accuracy plateaued below 0.4 after a small boost from the first to the second round.Figure 15A reports matcher accuracy with bootstrapped 95% confidence intervals and a chance level of 0.083.
- Practice results: Models perseverated on their previous response almost 100% of the time, with perseveration increasing slightly over repetitions.This pattern indicates increasing entrenchment in incorrect responses rather than adequate learning with practice.
- Interpretation: Invalid in-context information significantly worsened model performance when considered alongside the practice results and earlier findings.This conclusion connects the practice pattern with the broader interactive limited-feedback analysis.
- Response validity: Invalid responses from frontier models were relatively low overall and were ignored when calculating response probabilities and accuracies.Responses were invalid when models failed to return a parsable answer within 256 tokens.