Source-linked AI summary

The AI Assessment Scale (AIAS): A Framework for Ethical Integration of Generative AI in Educational Assessment

Mike Perkins, Leon Furze, Jasper Roe, Jason MacVaugh

arXiv:2312.07086v2cs.AI

TL;DR

Educational GenAI creates pedagogical opportunities alongside ethical and academic challenges, while its assessment implications remain underexplored. The paper proposes the AI Assessment Scale (AIAS), a graduated framework for clarifying and structuring GenAI use in assessment, and argues that it offers a practical institutional starting point for ethical integration. Its application includes boundaries: Level 3 is transitional, while Level 1 may raise equity concerns outside supervised or low-stakes settings.

  • Problem

    The ethical and pedagogical implications of integrating GenAI into assessment remain underexplored, despite growing educational use and unresolved concerns about integrity and equity.

  • Method

    The paper develops the AI Assessment Scale, a practical graduated framework that aligns permitted GenAI engagement with assessment purposes and clarifies expectations for educators and students.

  • Results

    The AIAS provides institutions with a flexible tool for clarifying GenAI use, adjusting assessments, supporting transparency, and maintaining academic-integrity-aligned practice.

  • Takeaways & Limitations

    The AIAS offers a practical starting point for shifting assessment policy from a binary AI/no-AI model toward nuanced, ethically structured GenAI integration.

  • Takeaways & Limitations

    Level 3 is described as a transitional stop-gap, while Level 1 may create equity concerns when unsupervised out-of-class work is permitted.

Abstract

from arXiv · show

Recent developments in Generative Artificial Intelligence (GenAI) have created a paradigm shift in multiple areas of society, and the use of these technologies is likely to become a defining feature of education in coming decades. GenAI offers transformative pedagogical opportunities, while simultaneously posing ethical and academic challenges. Against this backdrop, we outline a practical, simple, and sufficiently comprehensive tool to allow for the integration of GenAI tools into educational assessment: the AI Assessment Scale (AIAS). The AIAS empowers educators to select the appropriate level of GenAI usage in assessments based on the learning outcomes they seek to address. The AIAS offers greater clarity and transparency for students and educators, provides a fair and equitable policy tool for institutions to work with, and offers a nuanced approach which embraces the opportunities of GenAI while recognising that there are instances where such tools may not be pedagogically appropriate or necessary. By adopting a practical, flexible approach that can be implemented quickly, the AIAS can form a much-needed starting point to address the current uncertainty and anxiety regarding GenAI in education. As a secondary objective, we engage with the current literature and advocate for a refocused discourse on GenAI tools in education, one which foregrounds how technologies can help support and enhance teaching and learning, which contrasts with the current focus on GenAI as a facilitator of academic misconduct.

Introduction

GenAI is reshaping education by offering new pedagogical possibilities while creating ethical, academic-integrity, and critical-thinking challenges. The AI Assessment Scale (AIAS) responds with practical guidance for integrating GenAI into assessment.

  • GenAI models generate text, images, and audio from learned statistical patterns, and their increasing educational prevalence is prompting changes in teaching, assessment, and learning expectations.
  • Educational use of GenAI offers innovative teaching and learning possibilities but requires alignment with pedagogical objectives, academic integrity, ethical usage, and critical-thinking development.
  • Ethical and pedagogical implications of integrating GenAI into assessment remain underexplored, with gaps in both student–educator perspectives and cross-model integration research.
  • The AIAS provides clear expectations for student GenAI engagement, helps educators adjust assessments, and supports ethical assessment practices in higher education and potentially K–12 settings.

Literature

Research on GenAI in higher education documents substantial educational applications but presents an unsettled balance between benefits and harms. Student perspectives remain comparatively underexamined, and students may have limited experience and confidence using these tools.

  • GenAI can support complex-concept learning, communication accommodations, second-language learning, lesson planning, question generation, simulations, and role plays.
  • Whether GenAI’s benefits outweigh its drawbacks in higher education remains unsettled, with research and public discourse producing conflicting evaluations.
  • Student perspectives of GenAI in higher education have received little media and research attention compared with broader discussions of the technology.
  • Available student survey evidence indicates relatively low experience with GenAI and limited confidence in its learning and assessment applications.

Problematizing The View Of GenAI Content As Academic Misconduct

The paper challenges treating all GenAI-generated writing as academic misconduct and reframes the issue around educational integrity, cultural context, transparency, and responsible use. It proposes the AIAS as a standardised yet adaptable institutional response.

  • Academic-integrity discussions have been shaped by concerns about misconduct, cheating, and an arms race between technology-enabled dishonesty and detection software.
  • The authors argue that treating AI-generated writing as inherently incompatible with academic integrity is unsustainable for higher education’s future.
  • Publishers may permit or encourage GenAI for manuscript refinement when use is declared transparently and authors retain responsibility for accuracy and veracity.
  • Perceptions of plagiarism and academic-integrity violations are influenced by cultural values, while existing rules may not reflect diverse student populations.
  • The AIAS is proposed to support ethical GenAI engagement while giving institutions a standardised, adaptable approach to assessment policy.

The AI Assessment Scale

The AIAS developed as education moved from prohibiting GenAI toward a structured, graduated model of use. Its five-point scale balances simplicity with clarity and helps educators redesign assessments while clarifying ethical student use.

  • Development and Rationale: The AIAS emerged from a shift away from viewing GenAI primarily as plagiarism or misconduct and toward recognising its potential to enhance learning and performance.
  • Development and Rationale: The scale evolved from a binary AI/no-AI model into a more nuanced scaffold that allows discretion for teachers and students.
  • Development and Rationale: Its progressive design requires faculty to consider assessment restructuring while helping students use GenAI effectively and ethically.
  • Development and Rationale: The AIAS aims to help educators adjust assessments, clarify permitted GenAI use, and support academic-integrity-aligned submissions.
  • Scale Levels and Descriptions: The revised AIAS is presented as a five-point scale in Table 1.

Introduction to the Scale

The AIAS offers a flexible, cumulative scale that specifies permitted GenAI use and student responsibility across assessment tasks. It also recommends institutional policies and guidelines that address academic integrity, equity, and practical implementation.

  • The AIAS gives higher-education institutions a structured approach for specifying permitted GenAI use and student responsibility in assessments.
  • The scale is intended to be tailored by institutions rather than applied as a rigid linear model.Its simplicity is retained while accommodating the diverse nature of academic tasks and institutional policy decisions.
  • Each level is cumulative, so higher levels permit progressively broader forms of AI engagement.For example, Level 3 includes idea generation, structuring, and language editing, while Level 4 adds critical evaluation of AI contributions.
  • Level 1 prohibits GenAI use when assessments require students to rely solely on their own understanding, knowledge, or skills.Examples include technology-free discussions, in-class work, and viva-voce examinations.
  • Level 1 activities should be supervised or low-stakes because out-of-class no-AI conditions can create equity concerns.Students may differ in English proficiency, digital literacy, and access to advanced or paid GenAI tools.

Level 2: AI assisted idea generation and structuring

Level 2 permits GenAI for developing ideas and structuring work, while requiring the final submission to remain solely human-authored. It supports brainstorming, feedback, outlining, and research assistance without allowing directly generated content in the submission.

  • Level 2 permits GenAI for brainstorming, feedback, and structuring ideas, but the final submission cannot contain directly AI-generated content.
  • Students may use AI collaboratively to generate, discuss, filter, and refine ideas or to create structural outlines.
  • AI may provide research assistance by suggesting topics, areas of interest, or potentially useful sources.The passage specifies that source suggestions may use an Internet-connected model.
  • Level 3: Level 3 extends AI use to refining, editing, and enhancing students’ original language or content.Permitted applications include grammar correction, word-choice suggestions, structural rephrasing, and editing original images or videos.
  • Level 3: Level 3 requires students to submit original work alongside AI-assisted content to support comparison and authenticity.The authors describe this level as a transitional or stop-gap approach while assessments are more fully adapted to GenAI.

Level 4: AI Task Completion, Human Evaluation

Level 4 requires students to use GenAI for specified task components and critically evaluate the resulting outputs. Its flexible workflow accommodates iterative interaction between AI and human analysis, while Level 5 permits broader discretionary use.

  • Level 4: Level 4 requires students to critically assess AI outputs for relevance, accuracy, and appropriateness.The level emphasizes understanding both the capabilities and limitations of GenAI tools.
  • Level 4: Level 4 can involve direct AI generation followed by comparison with human-created content or development of an original response.
  • Level 4: Level 4 does not prescribe the sequence of AI and human work, allowing iterative rewriting after analysis when permitted.Any GenAI content must be cited appropriately, and deeper evaluation of AI-created content remains central.
  • Level 5: Level 5 allows AI use throughout an assessment at the student’s discretion or following teacher recommendations.Assessments may specify or recommend particular GenAI tools, or leave tool choice to students.
  • Level 5: Level 5 is suited to learning outcomes that require GenAI use or can be assessed regardless of AI usage.It also supports exploring GenAI as a collaborative and creative tool, including co-creation, experimentation, and continuous feedback.
  • A scalar approach is presented as necessary for clarifying academic-honesty boundaries across diverse digital-tool uses and supporting shared student-teacher understandings.

Balancing Skill Development, Engagement, and Ethics

The AIAS aims to balance academic integrity, skill development, and meaningful engagement by defining appropriate GenAI use and emphasizing ethical evaluation rather than misconduct alone. Its implementation must nevertheless address unequal access and the difficulty of keeping policies current.

  • The AIAS is intended to help institutions balance academic integrity, student skill development, and meaningful engagement with course content and assessment.
  • The scale promotes ethical GenAI use and develops students’ academic knowledge and tool-related skills as permitted AI use increases.
  • By defining acceptable AI-use parameters, the AIAS shifts attention toward skill development and ethical engagement rather than framing GenAI primarily as misconduct.
  • Unequal access to advanced or paid GenAI tools creates equity concerns when these technologies are integrated into assessment.Suggested responses include standardizing permitted tools or limiting access to certain advanced tools.
  • Keeping GenAI guidance practical is difficult because new tools continue to emerge.
  • The AIAS is primarily designed to help students understand how GenAI should be integrated into particular assessment tasks.The authors caution that misuse should not become the central focus of policy, although institutional rules and sanctions may still be needed.

Conclusion

The AIAS offers a customizable framework for clarifying acceptable GenAI use in assessment, while acknowledging limits in applicability, evidence, access, and enforcement.

  • The AIAS supports shared expectations for responsible and accurate GenAI use by educators and students.
  • The scale must be customized to programme or module learning outcomes and may not suit every assessment.
  • No empirical research yet establishes the AIAS’s effectiveness in reducing misconduct or improving student outcomes.
  • Because GenAI technologies and assessment practices evolve rapidly, the scale requires continuous maintenance, adaptation, and institutional dialogue.
  • The AIAS reframes GenAI guidance constructively through a five-point scale, rather than relying primarily on detection tools.
  • Access costs, location-based availability, and the digital divide constrain implementation across diverse educational settings.

Conflict of Interest

The authors report no actual or perceived conflicts of interest and no funding for the manuscript beyond university academic-time resources.

  • The authors disclose no actual or perceived conflicts of interest.
  • The manuscript received no funding beyond academic-time resourcing at the authors’ respective universities.
  • The authors used multiple ChatGPT modes based on GPT-4 to draft and revise text, then reviewed and took responsibility for the outputs.
  • An initial pre-peer-review manuscript version was posted on arXiv under a CC BY-NC-ND 4.0 licence.
Loading 2312.07086v2…