Source-linked AI summary

Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA

Nimol Thuon, Panhapin Theang

arXiv:2608.28635v1cs.CLcs.AI

TL;DR

Reliable understanding of low-resource Khmer documents remains uncertain despite advances in multimodal document models. This pilot study evaluates open Qwen-VL systems with direct, parser-assisted, and external OCR-assisted prompting, finding that external OCR helps while native Khmer-script understanding remains difficult.

  • Problem

    Khmer Document VQA is underexplored, and complex non-Latin script characteristics make reliable text recognition and document understanding difficult.

  • Method

    The study evaluates open Qwen-VL models on Khmer business documents using English and Khmer questions with original-form answers, comparing direct, parser-assisted, and external OCR-assisted prompting.

  • Results

    External OCR provides the strongest improvement, while larger Qwen-VL models improve direct performance but remain weak on Khmer-script and mixed-script answers.

  • Takeaways & Limitations

    Current MLLMs can handle some structured, English-visible, and numeric content, but reliable native Khmer document understanding remains an open challenge.

  • Takeaways & Limitations

    The pilot uses 62 document images and 370 question-answer pairs from one Khmer form-understanding collection, limiting coverage across document domains.

Abstract

from arXiv · show

Recent multimodal large language models (MLLMs) have advanced document understanding, visual question answering, and text extraction. However, their reliability in low-resource, non-Latin settings remains uncertain. Khmer form documents present particular challenges because they contain complex script forms, mixed Khmer-English fields, and monetary values in both Cambodian Riel and US Dollars. Available resources for Khmer Document VQA are also limited. This paper presents a pilot diagnostic evaluation of open MLLMs on Khmer document images. We construct an evaluation subset from the previously introduced KH-FUNSD collection, covering invoices, receipts, quotations, and other business forms. The subset includes questions in English and Khmer, with answers retained in their original English, Khmer, mixed-script, or numeric forms. Rather than introducing a full public benchmark, this study examines the capabilities and failure modes of existing models. We evaluate representative open Qwen-VL models using direct image-based prompting and compare parser-assisted and external OCR-assisted configurations with Qwen3-VL-8B. Direct Qwen3-VL-8B outperforms smaller models, achieving 51.9% overall accuracy, although performance remains limited for Khmer-script and mixed-script answers. External OCR produces the strongest results, reaching 61.9% with Tesseract and 61.6% with PaddleOCR. Nevertheless, Khmer-script answers remain substantially more difficult than English and numeric fields. The results indicate that current MLLMs can process visually clear English and structured numeric content, but reliable native Khmer document understanding remains an open challenge.

1 Introduction

Khmer Document VQA remains underexplored because native-script complexity, mixed-language fields, and currency variation challenge current MLLMs. This pilot study diagnoses model capabilities and failure modes rather than introducing a full benchmark.

  • Motivation: Khmer is a low-resource non-Latin language with complex stacked characters, dependent vowels, diacritics, dense compositions, and limited word separation.Real-world degradation from blur, low resolution, distortion, compression, uneven illumination, handwriting, and stamps further complicates recognition.
  • Motivation: Khmer business documents combine native Khmer, English, mixed-script fields, and Cambodian Riel or US Dollar values.These documents require simultaneous script recognition, field interpretation, and currency grounding.
  • Motivation: Current MLLMs can produce visually plausible Khmer document outputs while omitting, corrupting, or inventing linguistically invalid text.The contrast between plausible layouts and malformed content motivates diagnostic evaluation beyond visual appearance.
  • Research gap: Existing Document VQA benchmarks may overestimate generalization because they often emphasize high-resource languages, cleaner images, or less representative document collections.Khmer document understanding is consequently underexplored relative to mainstream benchmark settings.
  • Study scope: The paper conducts a Khmer-only pilot diagnostic study with English and Khmer questions while preserving English, Khmer, mixed-script, and numeric answers.The study examines whether open MLLMs answer selected Khmer business-document questions reliably and where they fail.
  • Study scope: The evaluation compares direct prompting, MLLM-parser-assisted prompting, and external OCR-assisted prompting while analyzing Khmer recognition, mixed-script confusion, currency grounding, and unsupported answers.Failure analysis focuses on the requirements for reliable native-script reading, field grounding, and calibrated OCR use.

2 Related Work

Related work establishes that Document VQA combines text reading, layout interpretation, and visual reasoning, while low-resource non-Latin settings remain inadequately tested. Khmer adds script, data, OCR, and mixed-language challenges beyond mainstream document benchmarks.

  • Document VQA: Document VQA extends general image VQA by requiring text reading, layout interpretation, table or form identification, and grounding in structured document content.Its difficulty arises from interaction among OCR-like reading, layout understanding, and visual reasoning.
  • MLLMs: Qwen-VL, Qwen2.5-VL, and Qwen3-VL represent recent progress in text-rich image understanding, document parsing, structured extraction, OCR, and multimodal reasoning.Qwen3-VL includes practical 4B and 8B variants for open-model evaluation.
  • Research gap: Strong performance on predominantly English or high-resource benchmarks may hide weaknesses on low-resource non-Latin documents.This motivates evaluating generalization beyond mainstream benchmark distributions.
  • OCR-assisted systems: OCR and document parsing pipelines provide recognized text, layout positions, and visual features that can support document question answering.Prior systems such as LayoutLM and LayoutLMv2 jointly model these document representations.
  • OCR-assisted systems: Parser-assisted prompting uses model-generated textual or structured evidence, whereas external OCR-assisted prompting supplies dedicated OCR output alongside the image and question.These settings test whether direct MLLM prompting is sufficient for low-resource scripts.
  • Khmer setting: Low-resource document understanding is constrained by scarce annotated data, OCR systems, layout benchmarks, and pretrained language resources.Non-Latin scripts add stacked forms, diacritics, shaping, irregular spacing, and degraded real-world imagery.
  • Khmer setting: Khmer Document VQA differs from prior Khmer OCR work because it requires question interpretation, field localization, answer reading, and preservation of script or numeric format.Prior Khmer studies primarily focus on OCR or scene-text recognition.
  • Khmer setting: Cambodian business documents commonly mix Khmer and English and may contain Cambodian Riel, US Dollars, or both.Invoices, receipts, and quotations therefore combine multilingual content with currency-rich fields.

3 Pilot Khmer Document VQA Dataset

The pilot subset uses real-world Khmer business forms to support controlled diagnostic analysis across language, answer type, question type, and difficulty. Answers retain their original script or format while normalization is deliberately conservative for Khmer.

  • Data source: The subset is drawn from KH-FUNSD and covers invoices, receipts, quotations, and other Khmer business or administrative forms.Documents include mixed Khmer–English text, structured fields, tables, company information, itemized rows, identifiers, dates, and currency values.
  • Data source: The study uses KH-FUNSD only for a small diagnostic subset and does not claim to introduce a full public Khmer Document VQA benchmark.Its purpose is to evaluate current MLLMs and identify recurring failure patterns.
  • Document coverage: Invoices, receipts, and quotations were selected because they contain recurring business concepts, structured fields, mixed-language content, and monetary values.They also represent practical Cambodian workflows involving extraction, search, and administrative processing.
  • Question design: Questions cover field extraction, name lookup, layout-dependent lookup, counting, and amount-related queries in English and Khmer.Khmer questions are translated or adapted from selected English questions while preserving the same answer target.
  • Answer design: Answers preserve their original English, Khmer, mixed-script, number, date, or currency forms, with raw and normalized fields retained.Normalization handles superficial formatting differences but avoids aggressive Khmer normalization.
  • Subset statistics: The pilot contains 62 document images and 370 question-answer pairs for controlled analysis across question language, answer group, question type, and difficulty.The subset is modest in scale and intentionally diagnostic rather than comprehensive.
  • Diagnostic labels: Difficulty labels function as diagnostic categories, with hard questions often requiring Khmer-script name recognition or native textual-field extraction.The question taxonomy reflects common business-document operations rather than a complete Document VQA taxonomy.

4 Experimental Setup

The experiments compare direct image prompting with model-generated parser evidence and external OCR evidence under a fixed short-answer JSON protocol. Evaluation uses normalized answer accuracy and separately measures output parsing.

  • Task: Khmer Document VQA is formulated as short-answer prediction from a document image and an English- or Khmer-language question.Answers may be English, Khmer, mixed-script, or number/date/currency strings.
  • Output protocol: Models are instructed to return a JSON object containing the predicted answer and a confidence label.They should preserve visible scripts and formats, returning Unknown when the answer is not visible or is uncertain.
  • Model settings: Direct prompting evaluates Qwen2.5-VL-3B, Qwen3-VL-4B, and Qwen3-VL-8B using only the document image and question.The setup targets practical diagnosis rather than a broad model leaderboard.
  • Assisted settings: The MLLM-parser setting first generates visible text, key fields, amount/date/currency fields, and layout notes before a second answering pass.The second pass receives the image, question, and generated parser text.
  • Assisted settings: External OCR settings provide Tesseract or PaddleOCR text together with the image and question while retaining image access.The same JSON format and preservation instructions apply across assisted settings.
  • Protocol boundaries: OCR engines are used without task-specific fine-tuning or manual transcript correction, and quantitative evaluation uses original images without hand-cropped evidence regions.These choices test ordinary parser or OCR assistance rather than a fully optimized Khmer OCR pipeline.
  • Metrics: Accuracy counts predictions matching raw or basically normalized answers, while Parse reports outputs following the expected JSON format.Normalization includes English lowercasing, whitespace cleanup, numeric comma removal, and simple currency aliases; unparseable outputs count as incorrect.

5 Results and Analysis

Direct Qwen3-VL-8B performs best among direct models, but native Khmer recognition remains a major weakness. External OCR improves overall and mixed-script performance, while uncertainty cautions against fine-grained OCR ranking.

  • Direct MLLM Performance: 51.9% overall accuracy: Qwen3-VL-8B achieves the best direct-prompting result, with 98.6% parsing success.Its gains are uneven across answer groups.
  • Direct MLLM Performance: 0.0% Khmer-script accuracy: all direct models fail on Khmer-script answers despite stronger performance on English and number/date/currency fields.This indicates that scaling alone does not solve native Khmer text recognition.
  • Direct MLLM Performance: Direct performance declines sharply on hard questions, which often require Khmer-script name recognition or native textual-field extraction.Name-question accuracy also remains low, whereas Qwen3-VL-8B substantially improves counting accuracy.
  • Parser and OCR-Assisted Evidence: 51.9% versus 51.1% overall accuracy: MLLM-parser assistance remains similar to direct prompting, while external OCR reaches 61.9% with Tesseract and 61.6% with PaddleOCR.The comparison covers direct, parser-assisted, and external OCR-assisted Qwen3-VL-8B configurations.
  • Parser and OCR-Assisted Evidence: 56.0% mixed-script accuracy: PaddleOCR raises performance from 14.0%, while Tesseract reaches 54.0%; Khmer-script accuracy reaches 16.0% with Tesseract but remains 0.0% with PaddleOCR.External OCR is more effective for structured and mixed-script evidence than for native Khmer answers.
  • Parser and OCR-Assisted Evidence: Wilson intervals show that small differences should not be overinterpreted, especially for 50-example answer groups.The direct mixed-script result is 14.0% with an interval of roughly 6.9–26.2%, compared with PaddleOCR’s 56.0% and roughly 42.3–68.8%.
  • Parser and OCR-Assisted Evidence: 11.4% to 37.1%: Tesseract produces the largest gain for name questions, while PaddleOCR raises total-amount accuracy from 46.6% to 61.6%.Name questions remain substantially harder than counting and field-value questions.
  • Failure Analysis: Representative failures include Khmer-script recognition errors, mixed-script confusion, layout-field confusion, currency grounding errors, OCR noise propagation, and plausible but unsupported answers.Examples include copying only the English portion of a mixed field and assigning a nearby amount to the wrong currency or total field.

6 Discussion and Limitations

The discussion finds that external OCR improves Khmer Document VQA, but native Khmer and mixed-script understanding remain unreliable. The pilot's limited coverage and evaluation protocol constrain how broadly its results should be interpreted.

  • 6.1 Discussion: External OCR provides the strongest improvement, especially for mixed-script and number/date/currency fields, yet Khmer-script answers remain much harder than English and numeric fields.The reported OCR-assisted results are strongest overall, while native Khmer remains the main difficulty.
  • 6.1 Discussion: Qwen3-VL-8B improves direct accuracy and output-format stability over smaller models but still misses many native-script fields.Its parsing reliability does not eliminate Khmer recognition errors.
  • 6.1 Discussion: MLLM-generated parser evidence provides little gain, suggesting that self-parsing often repeats the same recognition errors.External OCR is more useful than parser evidence generated by the MLLM itself.
  • 6.1 Discussion: Aggregate accuracy can hide severe failures in native-script fields, so evaluation must separately measure native-script and language-aware performance.The paper highlights names, native Khmer textual fields, and currency amounts as high-risk cases.
  • 6.2 Limitations and Future Work: The study covers 62 document images and 370 question-answer pairs from one Khmer form-understanding collection, limiting coverage across document domains.It focuses on open Qwen-VL models and two OCR engines.
  • 6.2 Limitations and Future Work: Normalized exact matching may mishandle equivalent Khmer text, names, or currency formats, and the study lacks oracle OCR transcripts to separate OCR from downstream errors.Confidence labels are collected but not yet calibrated.

7 Conclusion

The paper presents a pilot diagnostic study of open MLLMs for Khmer Document VQA, an underexplored low-resource non-Latin setting. Qwen3-VL-8B improves direct performance, while external OCR helps most, but robust native Khmer reading remains unresolved.

  • 7 Conclusion: The study evaluates open MLLMs for Khmer Document VQA, an underexplored low-resource non-Latin document setting.It examines current models rather than introducing a full benchmark.
  • 7 Conclusion: Qwen3-VL-8B improves direct performance and output-format stability over smaller models, but remains weak on Khmer-script and mixed-script answers.MLLM-generated parser evidence provides little improvement.
  • 7 Conclusion: External OCR assistance improves performance, especially for mixed-script and number/date/currency fields, although native Khmer-script answers remain difficult.The conclusion contrasts external OCR with self-parsing, which often repeats recognition errors.
  • 7 Conclusion: Current MLLMs are not yet robust Khmer document readers, with native-script recognition, mixed-script grounding, and currency-field interpretation remaining major challenges.They can handle some structured, English-visible, and numeric content.
  • 7 Conclusion: Future work should develop larger Khmer and multilingual non-Latin benchmarks, stronger Khmer-aware OCR and normalization, and evaluations targeting native-script and mixed-language understanding.These directions follow the paper's identified challenges and scope.
Loading 2608.28635v1…