Source-linked AI summary
Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts
Jaehee Kim, Ji Hoon Chung, Seoyoon Park, Unsol Kim, Kyungwon Park, Ji Hak Kim, Yi-Jun Chen, Hansaem Kim
TL;DR
Existing text-based and multimodal evaluations underrepresent indirect speech acts that require visual and sociopragmatic context, particularly in Korean. READI addresses this gap with a bilingual V-PQA benchmark using theory-grounded graded indirectness, and experiments show substantial difficulty as indirectness and contextual demands increase.
Problem
Prior studies and benchmarks largely use explicit textual context or perceptual recognition, leaving visually grounded, context-dependent pragmatic understanding underexplored.
Method
READI evaluates indirect directive interpretation through image–utterance V-PQA items with graded indirectness in English and Korean.
Results
Models struggle with visually grounded indirect speech acts; performance drops at higher intensities, with sharper degradation in Korean and errors primarily reflecting missed sociopragmatic connections.
Takeaways & Limitations
Multimodal integration alone is insufficient for recovering indirect intent, motivating evaluation of cultural and contextual pragmatic reasoning beyond visual–text alignment.
Takeaways & Limitations
READI uses theory-driven scenarios rather than naturally occurring conversations, limiting coverage of real-world spontaneity and variability.
Abstract
from arXiv · showhide
Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks largely overlook this requirement, focusing instead on explicitly encoded context or perceptual recognition, and thus underex- plore context-dependent pragmatic understand- ing, particularly in high-context languages such as Korean. We introduce READI, a multimodal benchmark for evaluating ISA understanding through integrated reasoning over visual con- text and dialogue. READI models graded in- directness grounded in pragmatic theory and formulates the task as vision-based pragmatic question answering (V-PQA), supporting cross- lingual evaluation in English and Korean. Ex- periments show that even state-of-the-art multi- modal models struggle with visually grounded indirect speech acts, with performance declin- ing as indirectness increases, underscoring the need for benchmarks that explicitly target con- textual pragmatic reasoning.
1 Introduction
Pragmatic understanding requires interpreting utterances through situational and sociopragmatic context, which is especially important for indirect directives. READI addresses this gap by evaluating whether models can infer directive intentions from visual context across English and Korean.
- Pragmatic meaning emerges from the relationship between an utterance and its situational context, not from linguistic form alone.
- Indirect speech acts require contextual inference because their intended directives are not directly expressed in the utterance.An utterance such as “It’s cold in here” can function as a directive to close a window.
- Prior LLM studies mainly present sociopragmatic factors as explicit text, which does not reflect conversational contexts conveyed through non-verbal and physical means.
- Existing multimodal evaluations emphasize image recognition, whereas this study focuses on pragmatics-centered contextual interpretation.
- READI formulates indirect speech understanding as V-PQA, combining interactional images, indirect directives, graded indirectness, and English–Korean comparison.
2 Related Work
Prior ISA and multimodal benchmarks inadequately represent visually grounded, pragmatically implied meaning. READI is motivated by the need to evaluate graded indirectness and context-sensitive interpretation, especially across languages and cultures.
- Indirect Speech Acts are utterances whose intended illocutionary force differs from their grammatical form or literal meaning.
- Indirectness does not function uniformly across cultures; Korean and Japanese deference is linked to interpersonal relations such as rank, status, and age.
- CCSARP distinguishes direct, conventionally indirect, and non-conventionally indirect directives by the amount of pragmatic inference they require.
- Text-centered ISA studies and existing VQA benchmarks omit non-verbal environmental cues and pragmatically implied meaning.
- Earlier multimodal HRI work found context-dependent directive meanings but largely tested conventionalized indirect expressions without graded indirectness or richly structured visual context.
- Benchmarks jointly covering indirect speech acts, graded linguistic indirectness, and visually grounded situational context remain scarce.
3 The READI Benchmark
READI is a theory-driven, bilingual benchmark in which models infer indirect directive intent by integrating utterances with image-grounded sociopragmatic context. Its controlled scenarios vary indirectness and contextual factors while using multiple-choice evaluation.
- READI evaluates whether models identify image-depicted sociopragmatic factors and use them to infer the intended meaning of indirect utterances.
- The benchmark uses CID, NCID-Strong Hint, and NCID-Mild/No Hint to represent increasingly demanding levels of indirectness.
- READI controls five interacting sociopragmatic factors: relations, power asymmetry, social distance, situation and setting, and directive target and content.
- The English and Korean subsets use language-specific scenario construction, native-speaker validation, and linguist review for pragmatic naturalness and label appropriateness.
- Images encode pragmatic context and exclude text or speech bubbles, so answers cannot be determined from linguistic information alone.
- The final benchmark contains 102 multimodal items: 57 Korean and 45 English, each pairing one image and indirect utterance with four directive-intent options.
- Given image I, utterance U, and options C, models estimate which option best represents the latent pragmatic meaning M.
- Multiple-choice options include literal traps and contextual mismatches to reduce surface-level guessing.
4 Experimental Settings
The experiments compare eight multimodal models across commercial, general-purpose open-source, and Korean-specific groups. Accuracy is used to compare their ability to identify intended pragmatic categories.
- Eight multimodal vision–language models were selected by commercial versus open-source status and Korean-specific training or adaptation.
- The models comprise commercial general-purpose references, 7–8B open-source general-purpose models, and Korean-specific multimodal models.
- READI uses accuracy as its primary metric because each balanced multiple-choice item has one correct intention category.
- F1 scores showed trends nearly identical to accuracy, so only accuracy is reported for clear cross-model comparison.
5 Results
READI results show that multimodal models struggle more as indirectness increases, with larger difficulties on Korean and under image–utterance mismatch. Language-specific adaptation helps especially under greater pragmatic complexity, but substantial errors remain.
- Overall performance: Most models perform worse on Korean than English, while Korean-specific models maintain or slightly improve Korean performance.The cross-language gap appears across commercial and open-source models; model scale or training-data size alone does not explain it.
- Performance by indirectness: English accuracy remains similar at Lv1 and Lv2 before sharply declining at Lv3, whereas Korean accuracy begins degrading at Lv2.Korean models show substantial difficulty at Lv3, where social relations, implicatures, and cultural expectations become important.
- Language adaptation: On Korean, language-adapted models outperform multilingual models as indirectness rises and show smaller performance drops in high-intensity conditions.Language adaptation contributes more to mitigating degradation under pragmatic complexity than to improving average accuracy.
- Image–context alignment: Shuffling images across items causes most models to lose accuracy on both datasets, supporting reliance on coherence between utterances and visual contexts.The result indicates that READI requires image-grounded contextual integration rather than inference from utterances alone.
- Qualitative error analysis: Models are strongest when utterances contain conventional markers or images directly indicate the target action, and errors increase when linguistic cues weaken.Observed errors include contextual misses, over- or under-interpretation, and cultural-context errors, especially in Korean.
6 Discussion
Multimodal models align images and text well but struggle to recover indirect speech-act intentions, especially in Korean. Errors reflect failures to connect sociopragmatic factors rather than object-recognition problems.
- Image-text alignment does not translate into reliable understanding of indirect speech-act intentions.
- Performance declines sharply at particular indirectness levels rather than gradually.English deteriorates at the highest intensity, whereas Korean declines from mid-level intensity.
- Korean performance is more affected because interpretation depends heavily on social and cultural context.
- Language-specific training mitigates performance collapse at high-intensity levels rather than merely raising average performance.
- The primary errors are contextual misses caused by failing to link sociopragmatic factors, not failures of object recognition.
7 Conclusion
READI evaluates multimodal understanding of indirect speech acts by integrating visual cues with pragmatic strategies. The findings show that multimodal integration alone is insufficient for recovering intent across linguistic and cultural contexts.
- READI is a multimodal benchmark designed to analyze models’ pragmatic inference capabilities stepwise.
- Multimodal integration alone is insufficient for recovering indirect-speech-act intent.
- The benchmark highlights performance gaps rooted in linguistic and cultural contexts.
8 Limitation
The study’s conclusions are limited by its narrow speech-act scope, theory-driven scenarios, two-language coverage, and focus on understanding rather than response generation.
- The analysis covers only indirect directive speech acts, excluding other types of indirect speech acts.
- READI uses theory-driven scenarios and expert review rather than naturally occurring conversations.This supports controlled manipulation of sociopragmatic variables but may reduce coverage of real-world spontaneity and variability.
- The benchmark includes only English and Korean, so its findings may not generalize readily to other linguistic and cultural settings.
- The evaluation measures understanding of indirect speech acts but not the generation of pragmatically appropriate responses.
A CCSARP Directness Levels
The cited materials identify the CCSARP directness-level table, annotation and inference guidelines, and sociopragmatic-factor table used in the benchmark documentation.
- Table 8 presents CCSARP directness levels with descriptions and examples.
- Table 9 presents detailed annotation guidelines and inference levels for indirect-speech-act data.
- Table 10 presents the sociopragmatic factors of the READI datasets.
C Task Examples
READI includes example instances in both English and Korean, presented separately in Tables 11 and 12.
- Table 11 presents an example from the English READI benchmark.
- The examples cover both English and Korean READI settings.
- Table 12 presents an example from the Korean READI benchmark.
D Detailed Performance Metrics for English and Korean READI
The English and Korean READI subsets are evaluated with Accuracy, Precision, Recall, and F1-score across graded indirectness levels and overall averages.
- The reported metrics are Accuracy, Precision, Recall, and F1-score.
- Performance is broken down across three indirectness levels and an overall average.
- Under the CCSARP framework, Level 1 denotes conventionally indirect directives, while Levels 2 and 3 denote non-conventionally indirect directives.
- Tables 13 and 14 report detailed performance metrics for the English and Korean READI benchmarks, respectively.