Source-linked AI summary

A Layered Taxonomy for Chinese Learner Grammatical Error Annotation

Mengyang Qiu, Jungyeul Park

arXiv:2609.02153v1cs.CL

TL;DR

The paper addresses a practical gap between computational schemes with stable labels and linguistically informative pedagogical analysis. It develops a layered taxonomy connecting edit-based CGEC annotation with linguistic and pedagogical analysis, and reports 24.0% orthographic errors plus extension triggers in another 17.6% of MuCGEC edits; the assessment supports complementary taxonomy layers.

  • Problem

    Computational schemes provide stable labels such as R:VERB but say relatively little about the grammatical meaning of learner errors, leaving a practical gap with pedagogical analysis.

  • Method

    The paper develops a layered taxonomy connecting edit-based CGEC annotation with linguistic and pedagogical analysis, using orthographic screening and linguistic error labels.

  • Results

    24.0% of examined MuCGEC edits were identified through orthographic screening, while an automatic procedure found extension triggers in another 17.6%; the assessment supported complementary taxonomy layers.

  • Takeaways & Limitations

    The taxonomy’s layers serve different but complementary purposes for interpreting Chinese learner errors.

  • Takeaways & Limitations

    The guidelines need clearer examples and decision rules, and reliability with trained human annotators remains to be established.

Abstract

from arXiv · show

Grammatical error annotation in Chinese learner writing requires labels that are both consistent and linguistically meaningful. This paper proposes a layered scheme linking computational Chinese grammatical error correction (CGEC) with pedagogical error analysis. The scheme first identifies character- and punctuation-level orthographic errors, labeling them by edit operation and subtype. Other errors receive a three-layer core label combining edit operation, linguistic domain, and part of speech, with optional Chinese-specific extensions for aspect, modality, comparison, argument structure, and complements. Drawing on CGEC resources, learner-error taxonomies, and Mandarin grammar, the taxonomy is evaluated through a coverage analysis of automatically extracted MuCGEC edits and a preliminary consistency study in which five large language models apply it to a sample. The results support the layered approach while identifying category boundaries requiring further refinement.

1 Introduction

The paper addresses a gap between stable computational error labels and the richer diagnoses needed for Chinese learner pedagogy. It proposes a layered taxonomy that preserves reproducible edit annotation while adding linguistically meaningful detail and evaluates it with MuCGEC coverage analysis and a five-LLM consistency study.

  • Motivation: Chinese learner writing poses annotation challenges because segmentation varies and grammatical meaning relies heavily on word order, particles, aspect, and constructions.Learner errors include homophone confusions, visually similar character substitutions, de-particle misuse, and 把/被 construction errors.
  • Research gap: Computational schemes offer stable labels but limited grammatical interpretation, whereas pedagogical classifications provide richer diagnoses without consistent corpus representations.The taxonomy is designed to bridge these complementary limitations.
  • Approach: The proposed scheme uses a two-route architecture: orthographic errors receive op:orth:subtype labels, while other errors receive op:dom:pos labels with optional extensions.The core layers encode edit operation, linguistic domain, and part of speech; extensions add finer functional and constructional distinctions.
  • Contributions: The study synthesizes CGEC resources and Chinese pedagogical taxonomies, then supplies category definitions, decision rules, and worked examples for modular adoption.The taxonomy is intended for manual, automatic, and LLM-assisted annotation.
  • Evaluation: The evaluation combines category coverage of automatically extracted MuCGEC edits with a preliminary annotation-consistency investigation involving five large language models.These tests examine empirical coverage and consistency of the proposed categories.

2 Prior resources, taxonomies, and design procedure

The design procedure draws together computational Chinese GEC resources, learner corpora, pedagogical taxonomies, and annotation principles. It retains distinctions that are observable and reproducible, places teaching-specific detail in extensions, and excludes interpretations requiring unsupported assumptions.

  • Design motivation: Word-based alignment can make segmentation choices alter the number and extent of extracted edits, motivating a representation that preserves edit information alongside linguistic interpretation.MuCGEC uses character-based edit spans and multiple reference corrections, while other resources distinguish minimal correction from fluency-oriented rewriting.
  • Source traditions: Chinese GEC resources support automatic edit extraction and surface-oriented labels, while learner corpora and pedagogical taxonomies provide more detailed information about learning difficulties and constructions.These traditions differ in segmentation, correction targets, linguistic categories, and pedagogical interpretation.
  • Procedure: The taxonomy was developed by comparing major CGEC datasets, learner-error resources, pedagogical classifications, and available annotation information.The comparison recorded annotation units, edit operations, segmentation, POS information, and Chinese-specific phenomena.
  • Selection criteria: Distinctions were retained when they were recoverable from the learner sentence, correction, and immediate context, applicable consistently across corpora, and independent of textbook sequence or feedback style.Orthographic errors receive operation and subtype labels; other errors receive operation, domain, and POS labels.
  • Selection criteria: Finer distinctions enter the extension layer when they aid Chinese teaching but are too detailed or context-dependent for cross-corpus comparison.Examples include ASP:le, ADD:dou, and CONST:ba.
  • Selection criteria: Distinctions requiring assumptions about internal causes, unexpressed learner intentions, or broader discourse pragmatics were excluded.This keeps annotation grounded in observable correction evidence rather than inferred causes.

3 A layered taxonomy for Chinese grammatical error anno-

The taxonomy treats an edit as the basic annotation unit and separates observable correction mechanics from linguistic and pedagogical interpretation. Orthographic errors follow a dedicated route, while other errors receive compositional domain-and-POS labels.

  • Annotation unit: An edit is defined as a difference between a learner sentence and one selected correction, with annotations kept separate when multiple corrections are possible.Projects may designate one primary correction or annotate each accepted correction separately.
  • Layered design: The layered design records how text changes, connects labels with Mandarin grammatical categories, and supports meaningful feedback.The layers separate correction mechanics from linguistic and pedagogical interpretation.
  • Edit operations: The operation inventory contains Missing, Replacement, Unnecessary, Word Order, and the orthographic-only Character Order operation.M inserts material, R exchanges forms, U deletes redundant material, WO marks word or phrase reordering, and CO marks internal character transposition.
  • Orthographic route: Orthographic errors receive op:orth:subtype labels for character spelling, character order, and punctuation phenomena.The scheme distinguishes phon, shape, complex, order, and punc subtypes.
  • Decision rules: Chinese-specific grammatical-function cases such as the de particles remain non-orthographic because their correction concerns grammatical function.The scheme therefore assigns them through the non-orthographic route rather than character-form labels.
  • Non-orthographic route: Non-orthographic errors receive op:dom:pos labels combining edit operation, linguistic domain, and part of speech.The domains are LEX-CONT, LEX-FUNC, and STRUCT, with POS tags including VERB, NOUN, ADJ, ADV, PART, AUX, PREP, CLF, and CONJ.

4 Pedagogical extensions

Pedagogical extensions add Chinese-specific functional and constructional detail to shared core labels without replacing them. Their modular design supports teaching-oriented refinement while preserving consistency and limiting ambiguity.

  • Extension purpose: Non-orthographic core labels provide a shared basis for corpus comparison and system evaluation, while optional extensions identify the relevant item or construction.Projects may adopt only modules matching their teaching or research aims, provided each adopted module is applied consistently.
  • Extension format: Extensions use family:item codes such as M:STRUCT:PART and ASP:le to encode both the core error and a finer pedagogical diagnosis.The extension is recorded separately from the core label.
  • Extension inventory: Word-level extension families cover temporal sequencing, additivity, contrastive stance, linking, location, aspect, and modality.These families group restricted sets of items sharing discourse, semantic, or grammatical functions.
  • Extension inventory: Constructional extensions cover focus, argument structure and voice, argument expression, comparison and degree, complements, and mixed constructions.Construction-level codes identify recurring grammatical patterns rather than isolated words.
  • Decision rules: Construction-based codes take precedence when a named construction organizes the edited material; otherwise, item-level codes are used.For example, a missing 都 in the 连· · · 都 construction receives FOC:lian-dou rather than ADD:dou.
  • Scope boundary: The extension layer is restricted to closed-class items and recurring constructions with small or clearly defined inventories.Open-class lexical errors remain at the core level unless projects document optional local subtypes.

5 Reference specification of the proposed taxonomy

The reference specification separates orthographic errors from non-orthographic errors, then assigns modular labels that combine edit information with linguistic and optional pedagogical detail.

  • Orthographic labels: Orthographic errors receive operation-based labels with sound-, shape-, order-, and punctuation-related subtypes, without requiring part-of-speech or pedagogical extensions.
  • Pedagogical extensions: Optional extensions encode finer-grained Chinese functional and constructional distinctions, including aspect, comparison, focus, argument structure, and complements.
  • Scope of examples: Examples are illustrative rather than exhaustive, and corrections dependent on intended meaning or discourse context state their interpretation in the English gloss.
  • Core labels: Non-orthographic errors receive a core label combining operation, domain, and part of speech, with X covering spans lacking a single applicable category.
  • Examples: The specification illustrates how missing, redundant, misplaced, or substituted elements receive distinct operations and domain-specific extensions.

6 Evaluation of the taxonomy

The taxonomy is evaluated for category coverage in MuCGEC edits and for preliminary labeling consistency among five LLMs, revealing broad coverage alongside boundary-sensitive disagreements.

  • Coverage analysis: 4,426 edits remain after excluding 58 sentences with annotator comments and 119 records that do not represent learner errors.
  • Coverage analysis: 55.9% of edits are Replacements, 27.4% Missing, and 16.7% Unnecessary; orthographic screening identifies 24.0% of all edits.
  • Coverage analysis: Orthographic edits divide almost equally between sound- or shape-based character substitutions and punctuation, while automatic reorder detection is rare.
  • Coverage analysis: 17.6% of edits fall into the extension inventory under a cautious function-word count, which misses multi-word repairs and shared-form content words.
  • Consistency study: Five LLMs labeled 391 items independently, with three models rerun to assess within-model consistency and agreement measured layer by layer.
  • Limitations: The study is preliminary because difficult divided cases are overrepresented, and reliability with trained human annotators remains unexamined.
  • Consistency study: Disagreements cluster around Replacement versus Missing, lexical-content versus lexical-functional analyses, X labeling, and extension applicability.

7 Discussion

The discussion presents the layered taxonomy as a bridge between correction operations and linguistic interpretation, while emphasizing unresolved reliability, coverage, and pedagogical-validation limits.

  • Interpretive structure: The layers serve complementary purposes: surface operations record changes, while domain, part of speech, and extensions represent increasingly specific interpretations.
  • Evidence and usability: The MuCGEC analysis suggests that dedicated orthographic and functional–constructional categories cover a substantial portion of the data, while surface labels are easier to apply than pedagogical interpretations.
  • Interpretive structure: Each annotation remains tied to a selected source–correction pair, preserving alternative target hypotheses instead of combining them into one diagnosis.
  • Practical integration: Compatibility with ChERRANT allows existing CGEC resources to reuse spans and operations while adding project-specific linguistic or pedagogical information.
  • Pedagogical scope: The extensions offer hypotheses about aspect, particles, comparison, argument structure, and complements that may support grammar-focused learner profiles.
  • Pedagogical scope: Their practical value remains to be tested with teachers and learners for interpretation, feedback, curriculum planning, and learner uptake.
  • Future validation: Broader validation should use multiple human annotators, corpora, learner backgrounds, proficiency levels, genres, correction policies, and edit-preparation methods.
  • Scope boundary: The scheme primarily addresses sentence-level changes recoverable from learner–correction comparisons; discourse phenomena may require wider context and additional layers.

8 Conclusion

The paper proposes a layered taxonomy separating observable edits from linguistic and pedagogical interpretation, linking CGEC annotation with Chinese learner-error analysis. MuCGEC coverage and preliminary LLM consistency results support the separation while highlighting the need for clearer guidance and broader validation.

  • Taxonomy architecture: The layered design connects edit-based CGEC annotation with linguistic and pedagogical analysis while preserving reproducibility.Observable edits are kept separate from their functional interpretation, supporting learner-corpus comparison and instructional analysis.
  • Taxonomy architecture: The taxonomy uses two routes: orthographic errors receive op:orth:subtype labels, while other errors receive op:dom:pos labels with optional extensions.Alternative accepted corrections remain tied to target hypotheses, allowing different analyses to be represented without collapsing them.
  • Coverage analysis: 24.0% of 4,426 MuCGEC edits were identified through orthographic screening, while a cautious automatic procedure found extension triggers in another 17.6%.These figures provide initial coverage evidence for separating orthographic and functional–constructional information.
  • Consistency study: Five LLMs applied the taxonomy to 391 items with greatest consistency for orthographic screening and surface edit operations.Agreement declined when part-of-speech and pedagogical distinctions were combined into complete labels.
  • Implications: The compact layers offer a comparatively stable basis for shared annotation, whereas more interpretive categories require clearer guidance.The conclusion treats this as preliminary evidence rather than a completed reliability evaluation.
  • Limitations and future work: Future work should test trained-human reliability, coverage across learner populations and genres, and whether extensions improve feedback and pedagogical interpretation.Discourse-level phenomena may also require additional layers because they cannot be diagnosed from a sentence–correction pair alone.

Declaration of generative AI in the manuscript preparation

During manuscript preparation, the authors used ChatGPT and Codex for language editing, restructuring, organization, and clarity. They reviewed the AI-assisted content, verified formal arguments, and accepted responsibility for the article.

  • Use of generative AI: ChatGPT and Codex assisted with language editing, sentence restructuring, manuscript organization, and technical clarity.The assistance concerned manuscript preparation rather than the paper’s reported taxonomy or evaluation.
  • Author oversight: The authors reviewed and edited all AI-assisted content.The declaration attributes final responsibility to the authors.
Loading 2609.02153v1…