VakyArth is the first pragmatic benchmark for Indic languages, covering Hindi, Punjabi, Tamil, and Malayalam. The benchmark diagnoses five pragmatic phenomena—deixis, speech acts, implicature, social pragmatics, and coherence—using multiple‑choice questions (MCQ), natural language inference (NLI), and translation tasks, all authored by native speakers.\ \ Experiments on multilingual large language models (LLMs) of various families and sizes reveal consistent failures when models must interpret meanings grounded in Indic linguistic and cultural conventions. Key observations include:\
- MCQ accuracy exceeds NLI accuracy for every model‑language pair;\
- Translation performance does not reliably track pragmatic understanding;\
- Indo‑Aryan languages enjoy a translation advantage over Dravidian languages.\ \ A deeper analysis shows that automatic translation metrics (e.g., BLEU, ROUGE) often miss fluent yet pragmatically unfaithful outputs, particularly for implicature and deixis.\ \ Review: VakyArth exposes systematic gaps in current multilingual LLMs’ pragmatic reasoning for Indic languages, highlighting the need for richer cultural and contextual signals in data, model design, and evaluation.