Source-linked AI summary
VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages
Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda, Anand Madhusoodanan, Geyang Guo, Wei Xu, Junyi Jessy L
TL;DR
Pragmatic evaluation has largely overlooked Indic languages despite their culturally diverse, context-dependent communication. VakyArth introduces a native-speaker-authored benchmark across four languages, five phenomena, and three task formats, finding systematic failures and task- and language-specific differences. The benchmark shows that literal defaults and automatic metrics can obscure pragmatic errors.
Problem
Existing pragmatic evaluation mainly targets English and high-resource languages, leaving Indic languages and their diverse pragmatic conventions without a dedicated evaluation.
Method
VakyArth evaluates Hindi, Punjabi, Tamil, and Malayalam across five pragmatic phenomena and MCQ, NLI, and translation tasks using native-speaker-authored items.
Results
MCQ accuracy exceeds NLI accuracy across all model-language combinations, while translation does not reliably track pragmatic understanding and automatic metrics miss some unfaithful outputs.
Takeaways & Limitations
Models’ pragmatic failures are systematic and often default to literal readings, making VakyArth diagnostic of errors that standard benchmarks and metrics may miss.
Takeaways & Limitations
The benchmark covers only single-turn utterances or short passages and four Indic languages, leaving multi-turn pragmatics and hundreds of languages uncovered.
Abstract
from arXiv · showhide
Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic languages unexplored despite their linguistic and cultural diversity. We introduce VakyArth, the first pragmatic benchmark for Indic languages, designed as a diagnostic evaluation covering Hindi, Punjabi, Tamil, and Malayalam. VakyArth evaluates models across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence; through multiple-choice questions, natural language inference, and translation, with all items authored by native speakers. Across multilingual large language models (LLMs) of varying families and sizes, we find consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions. Our analysis shows systematic differences across languages and tasks: MCQ accuracy exceeds NLI accuracy in all model-language combinations, translation performance does not reliably track pragmatic understanding, and Indo-Aryan languages show a translation advantage over Dravidian languages. We further show that automatic translation metrics can miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis.
1 Introduction
VakyArth addresses the lack of Indic pragmatic evaluation by testing context- and culturally grounded meaning across four languages, five phenomena, and three task formats. Results reveal systematic model failures, including literal interpretations and differences across tasks and language families.
- Motivation: Everyday Indic communication requires recovering meaning from context, cultural convention, and speaker intent rather than literal wording.The Hindi expression “beech wala” refers to the middle-priced suit, not the spatially or sequentially middle item.
- Research gap: Existing pragmatic evaluation focuses mainly on English and high-resource languages, while Indic pragmatics varies across linguistic families, scripts, and sociolinguistic communities.
- Contribution: VakyArth is the first Indic pragmatic benchmark, covering Hindi, Punjabi, Tamil, and Malayalam across deixis, speech acts, implicature, social pragmatics, and coherence.Items are native-speaker authored and include naturally code-mixed instances.
- Contribution: The benchmark evaluates pragmatic understanding through multiple-choice questions, natural language inference, and translation.These formats probe understanding at different levels of explicitness.
- Findings: Models show consistent failures rooted in Indic linguistic and cultural conventions, with MCQ accuracy exceeding NLI accuracy across all model-language combinations.Translation performance does not reliably track pragmatic understanding, and Indo-Aryan languages show a translation advantage over Dravidian languages.
- Findings: Automatic translation metrics can miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis.
2 Related Work
Prior pragmatic benchmarks cover limited phenomena and languages, with evaluation beyond English remaining restricted. VakyArth extends this work to pragmatic inference in Indic languages and culturally nuanced communication.
- Pragmatic evaluation of LLMs: Existing pragmatic resources are concentrated in high-resource languages and often cover a narrow set of phenomena.
- Multilingual and cross-cultural benchmarks: MultiPragEval evaluates English, German, Korean, and Chinese using Gricean maxims, while other recent benchmarks focus especially on Korean indirect speech acts.
- Indic language evaluation: Indic NLP benchmarks have expanded from foundational tasks to translation, summarization, and culturally specific knowledge, but pragmatic competence remains unaddressed.
- Indic language evaluation: PARIKSHA finds poor alignment between LLM evaluators and humans on culturally nuanced responses, while existing Indic benchmarks do not require pragmatic inference.VakyArth addresses implied meaning, deictic reference, and indirect speech acts.
3 Evaluation Design
VakyArth is designed around broad phenomenon coverage, task diversity, and cultural authenticity. Its items test context-dependent meaning through native-speaker-authored, often code-mixed examples across four Indic languages.
- Phenomena coverage: The benchmark spans five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence.These categories target non-literal, context-dependent meaning central to everyday communication.
- Deixis: Deixis items test context-dependent reference, including bidirectional temporal terms, proximal–distal contrasts, and multi-party spatial reference.Examples include kal and parso, Tamil intha–antha, and Punjabi edher.
- Speech acts: Speech-act items test indirect requests, commands, offers, refusals, and warnings whose intended actions diverge from surface form.Examples include a Hindi question functioning as a command and a Punjabi host’s ritual offer.
- Implicature: Implicature items require inferring meaning conveyed but not explicitly stated, often through culturally embedded expressions.Hindi hyperbole can express praise, while Tamil ritual indirectness can invite guests to stay and eat.
- Social pragmatics: Social-pragmatics items encode kinship, gender norms, power, face, and culturally appropriate forms of address.Examples include Punjabi mixed-gender greetings and Malayalam use of the kinship honorific chechi.
- Coherence: Coherence items require tracking discourse structure rather than selecting merely topically related continuations.Tamil meal-sequence examples distinguish a feast-completing response from comments that break the paragraph’s progression.
- Task formats: Each phenomenon is evaluated through MCQ, NLI, and translation formats that probe different levels of explicitness.MCQ distractors deliberately represent literal interpretations.
- Dataset construction: The dataset contains 578 items across Hindi, Punjabi, Tamil, and Malayalam, with native-speaker authorship, review, and removal of surface-solvable items.Dravidian languages require 1.6–2.5× more tokens per word than Indo-Aryan languages across three tokenizers.
4 Evaluation Setup
The evaluation compares five instruction-tuned multilingual models using three-shot prompts and Latin-transliterated inputs. Performance is measured with exact-match accuracy, COMET, and native-speaker human ratings.
- Prompting: All models and tasks use three-shot prompts with fixed same-language, same-script examples that are task-specific.Few-shot prompting was chosen because pilot zero-shot runs often produced unreliable output formats.
- Input format: All items are presented in Latin transliteration to reflect informal digital communication and reduce tokenization artifacts.The study also examines native-script versus Latin-script effects on pragmatic understanding.
- Metrics: MCQ and NLI use exact-match accuracy, while translation uses COMET and human 1–5 ratings for adequacy, fluency, and pragmatic adequacy.Human ratings come from four native-speaker raters, one per language.
5 Results and Analysis
VakyArth reveals recurring pragmatic failures across phenomena, tasks, languages, and models, including literal-reading biases, shared error cases, and translation-evaluation gaps.
- 5.1 Pragmatic Failure Patterns: Models frequently fail to resolve shifting temporal deixis, including Hindi and Punjabi references whose interpretation depends on discourse time.They mis-anchor bidirectional expressions such as “kal” and compute relative dates from the utterance rather than native conventions.
- 5.1 Pragmatic Failure Patterns: Models misinterpret culturally embedded expressions, indirect speech acts, kinship conventions, ritual politeness, sarcasm, and hyperbole across Indic languages.Examples include Tamil hospitality expressions, Hindi soft refusals, Punjabi care and sarcasm scripts, and Malayalam hyperbole.
- 5.2 Model and Task comparisons: Gemma-4-31B is strongest overall, scoring MCQ 0.86, NLI 0.43, and COMET 0.81, while Command-A-111B leads Gemma on NLI at 0.46 versus 0.43.Gemma remains most reliable despite being smaller than Sarvam-105B and lacking dedicated Indic pretraining.
- 5.2 Model and Task comparisons: MCQ accuracy exceeds NLI accuracy in all 20 of 20 model-language aggregates, indicating a consistent task-level performance gap.The benchmark distinguishes recognizing interpretations among options from inferring or generating them without scaffolding.
- 5.2 Model and Task comparisons: Tamil and Malayalam lag behind Hindi and Punjabi on translation, with the largest gaps in Coherence and Deixis and the smallest in Implicature.The authors attribute this likely translation advantage for Indo-Aryan languages to differences in pretraining representation rather than pragmatic difficulty.
- 5.3 Error patterns: 59.1% of wrong MCQ-answer events select the literal distractor, versus 33.3% expected by chance, while 24.7% of NLI items are failed by all four models.Shared failures reach 41.7% for Social Pragmatics, and automatic metrics are weakest for fluent but pragmatically unfaithful Implicature and Deixis translations.
- 5.5 Script Experiment: Native script produces 76.5% pooled MCQ accuracy versus 70.6% for Latin transliteration, but the effect varies by language and model.Punjabi with Qwen3-32B and Malayalam with Command-A perform slightly worse in native script.
6 Conclusion
VakyArth is presented as the first diagnostic benchmark for pragmatic competence in Indic languages, spanning four languages, five phenomena, and three task types. Its evaluation reveals systematic pragmatic failures, a persistent MCQ–NLI gap, and weaknesses in automatic translation metrics.
- VakyArth covers Hindi, Punjabi, Tamil, and Malayalam across five pragmatic phenomena and three task types.
- The benchmark finds phenomenon-specific failures that existing benchmarks cannot detect.
- Gemma-4-31B is the strongest model overall despite lacking Indic specialization.
- A persistent MCQ–NLI gap suggests that recognizing and reasoning about pragmatic meaning are distinct competencies.
- Automatic metrics can miss fluent but pragmatically unfaithful outputs for implicature and deixis.
7 Limitations
VakyArth has three stated limitations concerning the scope of its script experiment, its short single-turn items, and its language coverage.
- The script experiment covers only MCQ items, leaving fuller ablation across models, tasks, and languages for future work.
- The benchmark uses single-turn utterances or short passages and does not cover multi-turn pragmatics.
- Although the four languages represent diverse Indic language families, hundreds of other Indic languages remain uncovered.
A.1 Prompt Templates
The evaluation uses three-shot prompts, with instructions and examples written in each source language to avoid introducing English-language bias.
- All models and tasks use 3-shot prompts.
- Instructions and few-shot examples are written in the source language for each of the four languages.
- The prompt structure is identical across languages, while instruction text and examples are adapted by a native speaker.
Translation Prompt
The translation prompt asks models to translate Hindi text into English and output only the translated text. The supplied examples illustrate direct translations of Hindi utterances.
- The prompt instructs models to translate Hindi text into English and display only the translated text.
- The examples translate Hindi statements about working together, market goods, and keeping a promise into English.
- The template includes a placeholder for the source utterance.
MCQ Prompt
The benchmark materials illustrate how pragmatic interpretation depends on context, cultural convention, discourse structure, and speaker intent rather than literal wording alone. Across the examples, models fail on indirect meanings, reference, temporal anchoring, culturally conventional cues, and coherent continuation.
- MCQ Prompt: MCQ prompts require models to select an answer letter and its text from a contextualized Hindi question.The examples pair source contexts with four answer options and model-selected answers.
- NLI Prompt: NLI prompts ask models to classify relations between conversational premises and hypotheses as Entailment, Contradiction, or Neutral.The examples include food readiness, planned activity, and whether someone will go outside.
- A.2 Annotation Guidelines for Translation Rating: Human translation ratings assess preservation of implied meaning, tone, cultural nuance, and illocutionary force on a 1–5 scale.Annotators are instructed not to reward fluent outputs that miss pragmatic meaning or heavily penalize grammatical imperfections when meaning is preserved.
- A.3 COMET–Human Score Divergence Examples: COMET can assign high scores to translations that human raters judge as pragmatically poor because it rewards lexical and structural similarity.The paper presents cases with high automatic scores but human ratings of 1 or 2 out of 5.
- A.3 COMET–Human Score Divergence Examples: Translation failures include reversing Malayalam prohibitions, misassigning Punjabi insult targets, and misanchoring Punjabi temporal deixis.These outputs receive COMET scores of 0.78, 0.76, and 0.86 while humans rate them 1/5, 1/5, and 2/5, respectively.
- A.4 Error Analysis: The error analysis shows models missing a Tamil face-saving invitation and confusing family-dialogue referents across turns.No model recovers the implied invitation to stay, while Llama and Qwen conflate “amma” and “thangachi.”
- A.4 Error Analysis: Additional failures include reading Punjabi hyperbole literally, missing culturally caring criticism, and choosing topically related but discourse-incoherent Tamil continuations.Three models select plausible Tamil distractors that break the paragraph’s structured progression, while all four models miss the Punjabi cultural norm.