Source-linked AI summary
Dynamic Evidence Collection Ecosystem for Assessment Integrity and Authentic Competence
Rajan Kadel, Bellal Hossain, Samar Shailendra, Bushra Naeem
TL;DR
AI-generated final products weaken single-submission assessments as evidence of genuine mastery, while unreliable detection tools limit policing-based responses. This paper proposes a Dynamic Evidence Collection Ecosystem that captures continuous process evidence and reports improved learning visibility, academic integrity, and professional readiness.
Problem
AI-generated final products create a gap between correct submissions and genuine mastery, while single artefacts and unreliable detection tools weaken assessment validity.
Method
The paper proposes a Dynamic Evidence Collection Ecosystem using multimodal, multi-layered process evidence and AI-enabled assessment strategies integrated within the LMS.
Results
The proposed framework makes learning processes continuously observable, supporting judgements about student growth, decision-making, engagement, academic integrity, and professional readiness.
Takeaways & Limitations
The approach treats academic integrity as an integral part of assessment design by requiring sustained, multi-step, authentic engagement.
Takeaways & Limitations
This conceptual work lacks empirical validation, and its technological and governance requirements are not fully specified for scalability and adoption.
Abstract
from arXiv · showhide
Generative Artificial Intelligence (GenAI) can produce high-quality essays, code, and design artefacts, challenging the validity of conventional assessments that rely on single-point submissions and product-only grading. This paper proposes a design framework called "Dynamic Evidence Collection Ecosystem" that shifts assessment toward continuous, authentic, multi-source evidence of student learning over time. The framework collects process evidence through iterative artefacts, design logs, activity rounds, self-reflection, and peer collaboration, supported by an AI-enabled layer for learning analytics, formative feedback, and transparency. The approach is grounded in recent assessment-redesign scholarship in AI-rich contexts and aligned with contemporary views of authenticity in assessment. This paper builds on the hypothesis that academic integrity is strengthened when it is treated as an assessment design rather than as an AI detection problem. The tools have limitations and risks of use that carry academic penalties. This paper presents an implementation scenario to support institutional adoption.
I. INTRODUCTION
GenAI has made single-artifact submissions less reliable indicators of genuine competence, while AI-detection tools face reliability and ethical problems. The paper responds by proposing a Dynamic Evidence Collection Ecosystem and introducing 3-2-1 Portfolio Assessment as an example design.
- Assessment challenge: AI-detection tools show inconsistent performance and high false positives, particularly with improving model outputs and hybrid human–AI writing.False-accusation risks and non-transparent methods have led several universities to disable these tools.
- Assessment challenge: GenAI-generated final products widen the gap between correct submissions and genuine mastery of underlying skills.This creates a need to make the learning process visible to uphold academic integrity.
- Proposed contribution: The paper proposes a Dynamic Evidence Collection Ecosystem as a process-oriented response to GenAI-driven pressures on assessment.The framework positions assessment redesign as a way to make integrity a feature of design.
- Proposed contribution: The paper introduces 3-2-1 Portfolio Assessment as an example of assessment design.The paper then reviews relevant literature, presents the framework, demonstrates its application, and concludes the study.
II. RELATED WORKS
GenAI has challenged traditional product-based assessment centered on single final submissions, prompting institutions to use AI-detection tools despite evidence that these tools are unreliable for disciplinary decisions.
- Assessment challenges: GenAI has challenged traditional higher-education assessment by diminishing the worth of product-based evaluation through single final submissions.Conventional methods commonly assess essays, reports, or examinations as final products measuring student learning.
- Institutional responses: Higher-education institutions have responded to academic-integrity threats by adopting AI-detection tools including Turnitin AI, GPTZero, and ZeroGPT.
- Limits of detection: Peer-reviewed evidence finds AI-detection tools fundamentally unreliable and unsuitable as the primary basis for disciplinary outcomes.False positives are identified as a particularly serious concern, and current tools require major improvement.
III. DYNAMIC EVIDENCE COLLECTION ECOSYSTEM · A. Driver: Technology Push
The Dynamic Evidence Collection Ecosystem responds to GenAI-era assessment vulnerabilities by replacing single-submission, product-only assessment with a more authentic framework that improves learning visibility, academic integrity, and professional readiness. Its technology-push rationale centers on the devaluation of AI-generated artefacts, unreliable detection tools, and the hidden cognitive process in traditional assessment.
- III. DYNAMIC EVIDENCE COLLECTION ECOSYSTEM: Traditional single-submission assessments are increasingly vulnerable to AI misuse and provide limited visibility into how learning occurs.The literature review identifies both AI misuse and process invisibility as core weaknesses of conventional assessment.
- III. DYNAMIC EVIDENCE COLLECTION ECOSYSTEM: The proposed ecosystem offers a more authentic, robust alternative supporting learning visibility, academic integrity, and professional readiness.Fig. 1 presents the framework as a three-component model for authentic assessment.
- A. Driver: Technology Push: Assessment redesign is needed in the era of GenAI because product validity, AI-detection reliability, and process visibility are inadequate.These technology-driven factors motivate the component’s focus on changing assessment design.
- A. Driver: Technology Push: GenAI can generate essays, reports, designs, and code, weakening the validity of a single artefact as proof of learning.This devaluation creates a risk that students may misuse generated work.
- A. Driver: Technology Push: AI detection tools are unreliable, making policing AI use neither scalable nor pedagogically sound.The passage frames unreliable detection as a reason to redesign assessment rather than rely on policing.
- A. Driver: Technology Push: Traditional assessment observes outcomes without cognitive process, preventing distinction between student effort and AI output.The hidden-process problem leaves the relationship between learner effort and submitted work unclear.
B. Solution: Dynamic Evidence Collection Ecosystem · 1) Data-Driven Evidence:
The Dynamic Evidence Collection Ecosystem integrates multimodal data streams and AI-enabled assessment within the LMS. Its data-driven evidence layer builds a verifiable authorship trail through multiple inputs, including iterative artefacts, activity rounds, design logs, and prompting interactions.
- B. Solution: Dynamic Evidence Collection Ecosystem: The ecosystem combines multimodal data streams, AI-enabled assessment strategies, and seamless LMS integration.
- B. Solution: Dynamic Evidence Collection Ecosystem: It categorises collected evidence into two streams as part of its multimodal design.
- 1) Data-Driven Evidence:: The data-driven layer uses version-controlled artefacts and real-time learning analytics to create a verifiable trail of authorship and technical progression.
- 1) Data-Driven Evidence:: Iterative artefacts record incremental product development through drafts, prototypes, feedback cycles, and GitHub commits.They are especially suited to product design and development but can also support report writing.
- 1) Data-Driven Evidence:: Activity rounds provide longitudinal visibility through repeated cycles of proposals, drafts, feedback, revisions, improvements, and final artefacts.They may be particularly suitable for problem-based or project-based tasks.
- 1) Data-Driven Evidence:: Design logs document technical decisions, rationales, encountered problems, solutions, iterations, design changes, and analysed tools and resources.This evidence may be particularly suitable for design tasks.
- 1) Data-Driven Evidence:: Prompting interactions preserve student-AI dialogue histories, making AI-assisted workflows transparent by distinguishing student contributions from AI contributions.
2) Human-Driven Evidence:
Human-driven evidence adds human-in-the-loop verification of conceptual mastery through reflective journals, scaffolded peer feedback, and oral defence. These evidence streams should be evaluated together and synthesised through a centralised technological hub where AI acts as a facilitative partner.
- Human-driven evidence provides human-in-the-loop verification of conceptual mastery presented in the product.
- Reflective journals document students’ design choices and learning trajectories across their products, GenAI use, and other outcomes.
- Scaffolded peer feedback captures social learning, GenAI interaction, and students’ evaluation of their own and peers’ contributions in team environments.
- Oral defence enables real-time verification of authorship and conceptual understanding through spoken inquiry.
- Educators should evaluate human-driven evidence streams together, then synthesise them through a centralised technological hub with AI as a facilitative partner.
3) AI-Enabled Assessment:
The framework uses AI for process feedback, learning analytics, and AI-use disclosure to support assessment while preserving trust. These functions help educators interpret student progress and guide students before final submission.
- AI-Enabled Assessment: Together, these AI-enabled functions support educators’ interpretation of student progress without compromising trust.The framework combines process feedback, learning analytics, and AI-use disclosure within the assessment cycle.
- AI-Enabled Assessment: AI-use disclosure requires students to document how, where, and why AI supported their creative or technical process.This makes AI involvement explicit within the assessment process.
- AI-Enabled Assessment: Learning analytics automatically track engagement patterns to identify struggling students or work that deviates significantly from established norms.The analytics focus on changes in engagement and work patterns.
- AI-Enabled Assessment: AI-generated process feedback guides students during initial stages before final submission.The feedback is formative and supports students during the assessment cycle.
4) Technological Integration:
Technological integration centralises multimodal evidence, uses analytics to track learning over time, and supports feedback loops through an ePortfolio platform and LMS dashboard.
- Technological Integration: The ecosystem requires infrastructure that centralises multimodal evidence, aggregates it through analytics, tracks learning over time, and supports feedback loops.This integration manages high data volume while presenting outcomes in a suitable form.
- Technological Integration: An ePortfolio platform stores iterative artefacts, reflective journals, and video records as a longitudinal narrative of competence.Students curate these materials into a cohesive account of their development.
- Technological Integration: An LMS dashboard aggregates learning, assessment, design-log, and peer-interaction data into heat maps of class-wide and individual progress.Triangulating these multimodal streams supports verification of authentic competence and maintenance of academic integrity.
C. Outcome: Integrity and Authentic Competence
The dynamic evidence collection ecosystem strengthens integrity by making learning processes continuously observable and harder to fabricate than a single polished product. It also aligns assessment with professional practice through iteration, reflection, collaboration, and responsible tool use, while supporting discipline-specific implementation and feedback.
- Learning visibility: Continuous, contextual evidence makes learning processes visible and supports judgements about student growth, decision-making, and engagement.The framework provides observable traces that improve learning visibility, academic integrity, and professional readiness.
- Academic integrity: Ongoing, authentic engagement makes sustained, multi-step fabrication considerably more difficult than generating a single polished product.Integrity is embedded proactively in assessment design rather than addressed primarily through reactive detection.
- Professional readiness: The ecosystem mirrors professional practice through product iteration, self-reflection, collaboration, documentation, and responsible AI and tool use.These practices are intended to enhance employability skills, job readiness, and professional readiness.
- Implementation: Educators can adapt evidence selection and assessment strategies to their discipline, although the illustrated evidence mapping is primarily suited to STEM education.The paper recommends creating discipline-appropriate evidence tables covering purposes, applications, examples, and measured skills.
- Implementation: Educator reflection and evaluation feed improvements in evidence collection, AI-use strategies, and technological integration, shifting assessment from static products toward observing learning processes.The ecosystem uses feedback to support further development of its assessment design.
IV. 3-2-1 PORTFOLIO ASSESSMENT
The 3-2-1 Portfolio Assessment replaces a single major submission with staged, multimodal evidence collected across iterations, feedback, reflection, peer collaboration, and oral defence. It supports triangulated judgements of competence and AI transparency, but remains conceptual without empirical validation and with unspecified technological and governance requirements.
- Assessment design: Students document changes, feedback responses, decision reasons, and AI-tool use, including the tool, its purposes, and student modifications to outputs.The AI-use statement is intended to support transparency and value contextual reasoning and evaluative judgement.
- Assessment components: Design and project tasks require three stages or checkpoints that record decisions, constraints, progress, feedback responses, and changes from scope and planning through final product.Design tasks use decision and design ledgers; project tasks progress from scope and plan to prototype and final product.
- Assessment components: Report writing uses proposal, draft, and final versions, supplemented by two critique documents and one oral defence to assess reflection, peer review, authorship, and conceptual understanding.The oral defence enables real-time questioning of reasoning and independent articulation of the work.
- Implementation limitations: The framework requires educators to triangulate multimodal evidence, investigate missing or concerning components, and recognize that the conceptual design lacks empirical validation and fully specified technological and governance requirements.Future work should address these limitations through pilot testing.
V. CONCLUSION
GenAI challenges assessment systems built around single-point submissions and product-focused grading, while AI-detection technologies lack sufficient reliability for high-stakes integrity decisions. The conclusion therefore emphasizes assessment redesign centered on learning visibility, process-based evidence, and contemporary professional expectations.
- GenAI challenges the validity of assessment systems relying on single-point submissions and product-focused grading.
- AI-detection technologies lack the reliability required for high-stakes academic integrity decisions.
- Assessment redesign should foreground learning visibility, support integrity through process-based evidence, and align with contemporary professional expectations.