Source-linked AI summary
Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition
Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
TL;DR
Ancient Chinese artifact recognition lacks benchmarks with broad temporal, medium, script, and annotation coverage, despite its importance for heritage digitization. Ancient-Bench provides a 2,700-image benchmark with three tailored standardization protocols and evaluates general VLMs and OCR-specialist systems. Recognition remains fundamentally unsolved, while future extensions should add character-level spatial annotations and address possible data contamination.
Problem
Existing benchmarks are fragmented across temporal coverage, artifact media, script types, and annotation conventions, limiting consistent evaluation of ancient Chinese artifact recognition.
Method
Ancient-Bench combines 2,700 annotated images spanning 3,000 years, nine artifact categories, and seven script forms with symbol, character, and parsing standardization.
Results
Recognition remains fundamentally unsolved, with the best model achieving 52.08% NED and 57.86% F1 and particularly weak performance on oracle bones and bronze inscriptions.
Takeaways & Limitations
Ancient-Bench exposes persistent challenges in variant and rare characters, specialized symbols, layout-induced reading order errors, and hallucinations under visual ambiguity.
Takeaways & Limitations
The dataset currently provides line- or region-level rather than character-level spatial annotations, and publicly sourced data may have contaminated existing models.
Abstract
from arXiv · showhide
Ancient Chinese artifact text recognition is fundamental to heritage digitization, and benchmarks for ancient texts are essential for evaluating current model capabilities. However, existing benchmarks suffer from ''fragmentation'', manifested in limited temporal coverage, limited medium diversity, and incomplete script types. Therefore, we present Ancient-Bench, a comprehensive benchmark of 2,700 images for ancient Chinese artifact text recognition, featuring three dimensions: Multi-millennial (spanning 3,000 years of character evolution), Multi-medium (covering nine artifact categories), and Multi-script (encompassing seven historical script forms). To enable consistent and fair evaluation across heterogeneous media, we further define three annotation standards tailored to the medium-specific characteristics of ancient texts: symbol standardization, character standardization, and parsing standardization. Extensive experiments on Ancient-Bench covering general Vision-Language Models (VLMs) and OCR-specialist models reveal that ancient Chinese artifact text recognition remains fundamentally unsolved, with persistent challenges in variant characters, specialized symbols, and hallucination. The dataset is available at https://github.com/SCUT-DLVCLab/Ancient_Bench.
1 Introduction
Ancient Chinese artifact text recognition matters for cultural heritage preservation, yet existing benchmarks fragment evaluation across media, scripts, periods, and annotation conventions. Ancient-Bench addresses this gap with broad coverage, unified standards, and experiments showing that recognition remains fundamentally difficult.
- Digitizing ancient Chinese artifacts supports cultural heritage preservation and research on civilization’s evolution across more than three millennia.
- Existing benchmarks have limited medium, script, and temporal coverage and lack rigorous standards for symbols, variants, and structured layout parsing.
- Ancient-Bench contains 2,700 annotated images spanning 3,000 years, nine artifact categories, and seven historical script forms.
- The benchmark defines symbol, character, and parsing standardization protocols tailored to heterogeneous ancient artifact media.
- Experiments across general VLMs and OCR-specialist models find recognition fundamentally unsolved, with persistent difficulties involving variants, specialized symbols, and challenging inscriptions.Oracle bone and bronze inscriptions are particularly difficult.
2 Related work
Prior ancient-text benchmarks provide depth in selected media, scripts, or periods but remain fragmented. Their limited cross-dimensional coverage and inconsistent standards prevent systematic evaluation across eras, media, and scripts.
- Existing datasets remain limited in temporal span, medium diversity, script completeness, and systematic task definition.
- Single-medium datasets: Single-medium benchmarks range from printed historical books to multi-script documents, but remain constrained by selected media, scripts, periods, or artistic domains.
- Domain-specific datasets: Domain-specific datasets target oracle bones, bamboo or wooden slips, or calligraphy, providing specialized coverage without supporting broad cross-medium evaluation.
- Multi-medium dataset attempts: MCHDoc attempts multi-medium coverage but inherits inconsistent annotation standards and lacks unified specifications for artifact symbols, layout, reading order, and degradation.
- Benchmark fragmentation across medium, script, period, and task prevents recognition research across eras, media, and scripts.
3 Dataset Construction
Ancient-Bench is constructed from authoritative cultural-heritage sources using data-driven category formation and quality-oriented curation. The resulting benchmark spans nine media categories, seven script forms, three historical periods, and 2,700 annotated images.
- The construction process responds to fragmented datasets, inconsistent standards, and insufficient spatiotemporal coverage in ancient-text recognition.
- Ancient-Bench comprises 2,700 annotated images covering heterogeneous media, mainstream Chinese scripts, and more than 3,000 years of character evolution.
- Data Sources and Collection: Images and annotations were systematically collected and integrated from 14 national-level cultural heritage institutions under provenance and quality requirements.
- Data Sources and Collection: Sampling emphasized authoritative sources, diversity across media, scripts, and periods, and retention of representative naturally degraded samples.
- Dataset Taxonomy: The resulting taxonomy contains nine medium categories, seven mainstream Chinese script forms, and three historical periods.
4 Ancient-Bench
Ancient-Bench standardizes heterogeneous artifact data and annotations into a benchmark spanning diverse media, scripts, and historical periods. Its pipeline combines image preprocessing with symbol, character, and layout parsing standards, followed by multi-stage quality review.
- Dataset Scope: Ancient-Bench contains 2,700 annotated images covering nine medium types, seven script forms, and more than 3,000 years of history.The dataset draws on 14 national-level cultural heritage institutions.
- Data Challenges: Its source images vary in formats, regions, annotations, watermarks, artifact numbers, and background clutter, complicating fair recognition evaluation.Institutional annotations also differ in fidelity and symbol conventions.
- Data Processing Pipeline: A three-step preprocessing pipeline standardizes formats, filters prominent watermarks or irreversible damage, and localizes text regions for recognition.Silk manuscripts are exempt from damage-based filtering because degradation is inherent to the medium.
- Annotation Standards: The annotation framework defines symbol, character, and parsing standards to preserve medium-specific symbols, glyph forms, and original layout cues.Examples include retaining repetition marks and using placeholders for damaged or unrecognizable characters.
- Annotation Workflow: Annotation follows image preprocessing, symbol standardization, character standardization, and parsing standardization, with verification by trained annotators, experts, and an independent reviewer.The process took six months and involved 20 trained annotators.
5 Experiment
Experiments evaluate general VLMs and OCR-specialist models using character-level F1 and normalized edit distance. Results show low overall performance and medium-specific failures involving scripts, layouts, symbols, rare characters, and hallucinated outputs.
- Quantitative Experiments: 52.08% NED and 57.86% F1 are achieved by the best-performing model, Doubao-seed-2-0-lite-260428, while all models perform poorly overall.General VLMs outperform OCR-specific VLMs, and kimik2.6 achieves state-of-the-art performance.
- Medium-Specific Errors: F1 reaches only 15.75% for Gemini on Oracle and 38.48% for Doubao-seed-2-0-lite-260428 on Bronze, where models struggle with ancient and complex inscription glyphs.OCR-specific VLMs completely fail on both tasks.
- Medium-Specific Errors: Seal recognition yields 18.59% F1 for OCR-specific VLMs and 44.56% for general VLMs, with mirrored, relief, and hallucinated readings posing challenges.General VLMs also produce hallucinated outputs because seal inscriptions commonly contain classical poetry.
- Medium-Specific Errors: Slip, silk, cliff, edition, and calligraphy recognition expose confusion among similar or rare characters, special-symbol errors, degradation effects, and reading-order failures.Silk recognition includes hallucinated outputs in damaged regions, while cliff and edition tasks present distinct reading-order challenges.
- Hallucination Errors: Calligraphy hallucinations arise from prior-driven semantic completion and inaccurate cropping or detection that introduces neighboring strokes or noise.These mechanisms produce substitutions, insertions, deletions, spurious characters, and concatenated readings.
6 Conclusion
Ancient-Bench is introduced as a broad benchmark for ancient Chinese artifact text recognition, combining multi-millennial, multi-medium, and multi-script coverage with unified annotation standards. Evaluations show that the task remains unsolved, with weak performance and recurring recognition failures.
- 2,700 annotated images span 3,000 years, 9 media types, and 7 historical script forms.
- Three unified standards address symbol standardization, character standardization, and parsing standardization.
- Failure modes include variant or rare character confusion, missed symbols, layout-induced reading-order errors, and hallucinations under visual ambiguity.
Limitations
The benchmark’s current annotations support end-to-end recognition evaluation but leave character-level spatial detail for future improvement. Publicly sourced data also create a potential contamination concern.
- Current annotations provide line-level or region-level transcriptions for evaluating end-to-end recognition models and MLLMs.
- Character-level bounding boxes are planned to support fine-grained text detection evaluation in future updates.
- Because the data come from publicly available repositories, training contamination by open-source or closed-source models cannot be guaranteed absent.
Ethical Statement
Ancient-Bench uses images published by cultural heritage institutions and other publicly accessible sources, while restricting the benchmark to academic research use. Users remain responsible for following applicable licenses and terms.
- Images come from major cultural heritage institutions and other publicly accessible sources.
- The project does not claim or transfer copyright in the images.
- Users must comply with applicable licenses and terms of use, and the benchmark is intended only for academic research.
A.1 Medium and Chinese Script Categories
Ancient-Bench organizes ancient Chinese artifact texts across a continuous trajectory of more than 3,000 years and three historical periods. The listed media exhibit distinct scripts, material conditions, and layout conventions.
- Temporal organization: The benchmark’s temporal trajectory spans over 3,000 years and is organized into three historical periods.
- Early Civilization Period: Oracle bones cover the Late Shang to early Western Zhou, featuring pictographic incisions, surface cracks, variant characters, and ligatures.
- Cliff Inscriptions: Cliff inscriptions span Han to Tang–Song periods and suffer blurred characters and fragmented strokes from weathering, erosion, and rock textures.
- Early Modern Printing and Artistic Maturity Period: Woodblock editions use mature layouts such as centerfold strips and interlinear notes, with seal, clerical, regular, cursive, and running scripts.
- Early Modern Printing and Artistic Maturity Period: Calligraphy materials center on Ming and Qing ink-on-paper works, including correspondence, poetry drafts, and painting or calligraphy inscriptions.
A.3 Professional Paleographic Reference Tools
Ancient-Bench uses professional paleographic tools and multi-stage verification to support medium-aware annotation and evaluation. Its prompts and normalization rules preserve ancient text characteristics while standardizing outputs for fair comparison.
- Reference tools: Professional reference platforms provide glyph databases, transcriptions, and search tools for locating and comparing ancient Chinese characters.Yinqi Wenyuan integrates oracle-bone rubbings, glyph forms, and literature; Guyin Xiaojing supports glyph lookup and comparison.
- Annotation workflow: The annotation pipeline cross-validates artifact images against transcriptions and specialized glyph databases before multiple annotator reviews and final approval.Oracle-bone examples are checked character by character, then validated with Guyin Xiaojing in a second review round.
- Prompt design: The base rules require outputs to preserve visible traditional, rare, and variant characters, follow actual reading order, and avoid completing unseen text.Outputs also use prescribed placeholders and formatting tags, including <res> and <unrecognizable>.
- Prompt design: Task-specific prompts address document-specific structures because a single unified prompt is insufficient across heterogeneous document types.Woodblock editions, for example, contain interlinear notes and page margins that require specialized guidance.
- Evaluation normalization: Evaluation normalization treats special tags and repeated placeholders as atomic or count-matched tokens while removing whitespace and retaining simplified–traditional distinctions.The procedure aims to reduce penalties caused by formatting noise without erasing genuine character-form differences.