Source-linked AI summary

Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment

Jahangir Alam SM, Md Khalid Syfullah, Saad Ahmed, Munira Akter Mou, A K Z Rasel Rahman, A. K. M. Masudur Rahman, Mohammed Sowket Ali

arXiv:2608.22346v1cs.CVcs.AI

TL;DR

Research on rubric-aware examination assessment needs multimodal answers linked to question context and criterion-level expectations, not only images and total marks. This article constructs and validates such a dataset, comprising 485 submissions with structured OBE metadata, and confirms record, score, and file integrity across the collection.

  • Problem

    Rubric-aware assessment datasets need criterion definitions, performance descriptions, criterion marks, questions, and reference answers alongside examination images and total marks.

  • Method

    The authors consolidated scanned answers and expert assessment files from four institutions into randomized PDF-JSON pairs with question-linked OBE rubric metadata.

  • Results

    The final dataset contains 485 answer submissions, 12 question templates, and 47 rubric criteria, with totals validated against criterion marks and rubric maxima.

  • Takeaways & Limitations

    The collection supports research on multimodal document understanding, rubric-aware evaluation, criterion-level feedback, score prediction, and auditability under varied acquisition conditions.

  • Takeaways & Limitations

    The dataset covers only nine subjects and 12 question templates with uneven subject distribution, so cross-question comparisons require normalization or question-aware modeling.

Abstract

from arXiv · show

This data article describes a multimodal collection of scanned examination answers paired with expert-designed Outcome-Based Education (OBE) grading metadata. The collection contains 485 answer submissions from 415 consenting students at four academic institutions. Eight faculty contributors supplied examination materials covering nine subjects and 12 distinct question templates. Each answer-level item links a scanned PDF to a randomized identifier, subject label, question, model answer, criterion definitions, performance-level descriptions, criterion marks, and a total mark. The 12 rubrics contain 47 criteria in total. The scans retain realistic academic content, including handwriting, printed text, equations, tables, code, figures, sketches, and diagrams. CamScanner, Adobe Scan, and conventional scanners contributed variation in illumination, contrast, orientation, compression, and resolution. Diverse handwriting, crossed-out work, revised calculations, and inserted corrections add further visual variability for robustness and generalization studies. Preparation involved heterogeneous-source consolidation, label and text standardization, score validation, identifier randomization, filename randomization, and JSON-to-PDF integrity checks. An answer-level audit confirmed 485 unique identifiers, 485 unique PDF filenames, agreement between each total mark and its criterion-mark sum, and scores within the applicable rubric maximum. The data can support rubric-aware automated evaluation, multimodal document understanding, criterion-level feedback, score prediction, and privacy-aware OBE assessment research. Access is restricted to research use and is available from the corresponding author upon reasonable request.

1. Value of the Data

The collection links scanned examination evidence with question context and expert-authored OBE grading metadata at answer level. Its diverse subjects, visual conditions, and criterion-level labels support robust multimodal understanding, rubric-aware assessment, and auditable feedback research.

  • Answer-level records connect visual examination evidence with questions, model answers, OBE criteria, performance descriptions, criterion marks, and total marks.
  • Nine-subject coverage and mixed handwritten and printed content support document understanding, optical character recognition, vision-language modeling, and cross-subject generalization.
  • CamScanner, Adobe Scan, and conventional scanner inputs introduce realistic variation for robustness studies across lighting, handwriting, corrections, page geometry, and capture conditions.
  • Criterion-level labels enable systems to identify which answer components satisfy examiner expectations rather than reporting only an aggregate score.
  • The paired PDF-JSON structure supports experiments in score prediction, rubric selection, evidence grounding, feedback generation, calibration, and auditability.

2. Background

Outcome-Based Education links assessment to explicit learning evidence, criteria, and marks, while analytic rubrics preserve criterion-level performance information that overall scores omit. This collection pairs complete scanned answer submissions with structured rubric-aware records for transparent research, not official grading or model-performance comparisons.

  • Assessment foundations: OBE connects each question and response to intended learning evidence, examiner criteria, and assigned marks.Overall scores alone do not show which required components were demonstrated, omitted, or partially expressed.
  • Assessment foundations: Analytic rubrics divide tasks into criteria and describe performance across multiple quality levels.Rubric-aware data must retain criterion definitions, performance descriptions, maximum allocations, criterion marks, and links to the question and reference answer.
  • Document understanding: Real examination scripts expand automatic grading into document understanding across handwriting, equations, tables, code, graphs, diagrams, layout, and visual evidence.Document-AI systems therefore require data preserving both page appearance and structured assessment context.
  • Dataset contribution: Each complete answer submission pairs a scanned PDF with answer-level JSON containing subject, question, model answer, total mark, and rubric criteria with marks and performance rules.This organization supports tracing predicted scores to criterion-specific evidence.
  • Scope and limitation: The article addresses data organization and preparation rather than model comparisons or operational grading performance.The dataset is intended for controlled research and not for determining official student grades.

3. Data Description

The dataset comprises 485 answer submissions from 415 students across four institutions, with scanned responses paired to structured OBE grading records. It spans diverse academic content, acquisition conditions, question templates, and criterion-level rubrics for multimodal assessment research.

  • Collection scope: 485 answer submissions came from 415 participating students across four institutions, with materials supplied by eight faculty contributors.The answer-level organization allows one participant to contribute more than one submission.
  • Visual content and variability: Scans preserve text, equations, tables, figures, diagrams, code, mixed layouts, handwriting variation, and authentic corrections across varied acquisition conditions.CamScanner, Adobe Scan, and conventional scanners introduced differences in lighting, contrast, orientation, cropping, compression, and resolution.
  • Scope and limitation: The dataset broadens visual-domain coverage for generalization studies, but external validation is needed beyond the participating institutions and subjects.The stated scope concerns writing styles, document content, and acquisition conditions represented in the collection.
  • Data organization: Each answer pairs one randomized PDF filename with one JSON record containing the question, model answer, total mark, rubric criteria, gained marks, and performance descriptions.The answer-level structure keeps grading context next to the file reference while leaving the PDF unchanged as visual evidence.
  • Score representation: Subject-level mean marks are normalized by each answer’s question-rubric maximum, making questions with 5, 6, 9, 10, or 16 available marks comparable.The resulting mean normalized mark is descriptive metadata rather than a model evaluation.
  • Rubric structure: 12 question templates contain 47 criteria, with rubric maxima ranging from 5 to 16 marks and question-specific criterion allocations.Nine questions use four criteria, two use three criteria, and one uses five criteria.

4. Experimental Design, Materials and Methods

The dataset was assembled from consented examination scripts and expert-designed OBE grading metadata, then standardized into linked answer-level PDF-JSON records. Preparation preserved multimodal variability, validated marks and mappings, and supported privacy-aware research use.

  • Data collection: 485 answer submissions from 415 consenting students at four institutions were supplied by eight faculty contributors.Contributors provided scanned scripts and examiner-prepared assessment spreadsheets.
  • Multimodal variability: Preserved scans retain handwriting, equations, tables, code, figures, corrections, and acquisition variation from mobile and conventional scanners.Variations include illumination, contrast, shadows, orientation, cropping, compression, resolution, handwriting style, and correction density, supporting grouped robustness evaluation.
  • Data structure: Each answer record links a scanned PDF with the subject, question, model answer, OBE rubric, performance descriptions, criterion marks, and total mark.Rubrics were question-specific and retained variable criterion counts, including three- and five-criterion examples.
  • Validation: The final audit confirmed that every total equals its criterion-mark sum and that no total exceeds its applicable question-level rubric maximum.Question 1 scores ranged from 0 to 5, with a mean of 2.82 marks.
  • De-identification and integrity: 485 unique IDs and 485 unique randomized PDF filenames were verified with one PDF reference per record and preserved record-to-file mappings.The preparation process also checked that each referenced PDF existed in the assembled directory.
  • Preparation workflow: 485 records were consolidated, standardized, validated, and organized as one-to-one answer-level PDF-JSON pairs.Processing addressed heterogeneous file naming, spreadsheet organization, question formatting, criterion formatting, labels, and mark representations.
  • Privacy and reuse: Access is restricted because some source images may contain residual identifiers or educational records.Record-level de-identification does not guarantee removal of embedded text or logos.

5. Limitations

The collection’s reuse is constrained by limited and uneven coverage across subjects and question templates, alongside question-specific rubric differences that complicate direct comparisons.

  • Coverage and comparability: 9 subjects and 12 question templates limit the collection’s coverage, with an uneven subject distribution.E-commerce has 5 answers, whereas Machine Learning has 88.
  • Coverage and comparability: Question-specific rubrics vary in maxima, criterion counts, and textual scales, requiring normalization or question-aware modeling for comparisons.These differences make direct cross-question comparisons inappropriate without adjustment.

6. Ethics and Privacy Statement

The examination scripts were prepared and used for research with student consent, independently of official grading decisions. Access is restricted to approved research purposes, with privacy safeguards and human oversight required for any automated assessment use.

  • Consent and grading independence: Student consent covered preparation and research use, while the dataset and activities remained separate from official grading decisions.The data did not determine, alter, or replace participating students’ direct grades.
  • Privacy safeguards: Answer IDs and PDF filenames were randomized, but some page images retain names, student IDs, university names, logos, or related identifiers.These identifiers remain only within the controlled research context.
  • Privacy safeguards: Access is limited to approved research purposes, requiring recipients to protect files, avoid re-identification and public redistribution, and report only aggregate or de-identified information.These requirements govern handling and reporting of the dataset.
  • Automated assessment governance: Future automated assessment use must retain human monitoring, and research-model outputs cannot serve as official grades or the sole basis for decisions affecting students.Human oversight remains mandatory when automated assessment is used in future research.

Declaration of Generative AI and AI-Assisted Technologies in the Manuscript Preparation Process

Generative artificial intelligence assisted manuscript preparation, while the human authors reviewed the work and assumed full responsibility for its final content.

  • Generative artificial intelligence was used to assist with manuscript preparation.
  • The human authors monitored and reviewed the manuscript, numerical statements, tables, figures, and references.
  • The authors take full responsibility for the final content.
Loading 2608.22346v1…