Source-linked AI summary

A Survey Instrument to Assess Students' AI and Generative AI Knowledge

Aditya Johri, Cory Brozina, Akriti Bagale

arXiv:2608.21391v1cs.CYcs.AI

TL;DR

Existing assessments often do not measure technical concepts, practical applications, and ethical concerns together, motivating a broader objective instrument. The paper curates and implements a multi-source survey for higher education, finding useful diagnostic capability alongside substantial misunderstandings of AI mechanisms and higher-level concepts.

  • Problem

    Existing AI assessments often focus on attitudes or limited domains, leaving a need for an efficient instrument that measures factual knowledge across technical, practical, and ethical aspects of AI and GenAI.

  • Method

    The authors curate a 30-item survey from multiple instruments, align it with four AI-literacy aspects, and implement it with undergraduate IT students.

  • Results

    The instrument was comprehensive and discerning, identifying common misconceptions despite students’ generally good understanding of where AI is used.

  • Takeaways & Limitations

    The survey can support formative or diagnostic curriculum design by identifying knowledge gaps, including students’ need for deeper understanding of AI architectural principles.

  • Takeaways & Limitations

    The study used a single university and a single student cohort, so sampling bias may affect responses.

Abstract

from arXiv · show

In this research-to-practice paper we present a survey that can be used to assess students' AI knowledge. As the use of artificial intelligence (AI), including generative artificial intelligence (GenAI), has proliferated, so has the need to educate students about the topic. A range of AI literacy frameworks have been proposed, outlining the essential knowledge that students should have. Alongside, different ways of assessing AI knowledge have been developed. As yet, there is a lack of assessment instruments capable of evaluating multiple forms of student knowledge, including technical concepts, practical applications, and ethical concerns about AI use. In this article, we present a study implementing a comprehensive instrument to assess AI knowledge. The instrument combines measures from multiple scales to capture a range of literacy features and actual knowledge. We implemented the instrument in a higher education setting to assess its viability and usefulness and found that the instrument exhibited useful diagnostic capabilities and was able to identify common misconceptions among students. Although students performed well overall, there was a significant misunderstanding of how AI, especially GenAI systems, work. It also identified a lack of higher-level knowledge. The instrument is publicly available for use by others. We foresee its usefulness as a diagnostic that goes beyond understanding students' attitudes and perceptions of AI and GenAI use and tests multiple aspects of students' knowledge and conceptual understanding. This can enable the development of targeted instruction.

I. INTRODUCTION

The paper addresses the need to educate students about AI and assess their knowledge across technical, practical, evaluative, and ethical dimensions. It presents a survey curated from multiple sources to provide a broader assessment of AI literacy.

  • Motivation: AI’s expanding use in education increases the need to educate students and assess what they know about the technology.Higher education institutions are developing AI guidance while seeking ways to improve students’ AI expertise.
  • Assessment gap: Existing AI-literacy assessments often remain exploratory or lack comprehensive coverage of students’ knowledge.The paper notes that hands-on assessments can be situationally appropriate, but short surveys remain useful for this purpose.
  • Survey scope: The proposed survey covers technical knowledge, AI use, and ethical and responsible aspects of AI.Its items are intended to examine multiple elements of AI literacy rather than a single knowledge dimension.
  • AI literacy constructs: AI literacy includes knowing fundamental concepts, using AI tools, critically evaluating outputs, and navigating ethical and societal implications.The framework includes understanding how applications work, accomplishing tasks with AI, interpreting outcomes, and considering fairness, accountability, privacy, misinformation, and bias.

B. AI Literacy Assessments

The paper reviews existing AI-literacy instruments and identifies a need for an efficient objective assessment covering both AI and GenAI knowledge. It therefore curates items from multiple validated or established instruments.

  • Assessment gap: Many existing instruments assess AI attitudes or experiences rather than factual knowledge aligned with AI-literacy constructs.The paper focuses on objective knowledge while acknowledging that prior instruments often target perceptions or intervention efficacy.
  • Curation strategy: The curated survey combines items from multiple instruments to efficiently measure actual knowledge across all AI-literacy elements.The authors modified items to suit their purpose and aimed to capture a broad scope within a short survey.
  • Scope: The survey covers both basic and advanced knowledge across AI and GenAI, whereas other surveys focus on only one of these domains.This broader scope is presented as a distinguishing feature of the curated instrument.
  • Related instruments: GLAT is a 20-item multiple-choice assessment of AI and GenAI knowledge whose scores predicted performance on GenAI-supported tasks.GLAT was validated with 355 higher education students and outperformed self-reported proficiency measures.
  • Curation strategy: The authors used most GLAT items with modifications, relocating broader AI items into the AI portion and adding items about AI generally.The resulting survey adapted GLAT rather than developing every item from scratch.

B. Pew Survey

The Pew-derived material supplies survey items spanning AI recognition, application, evaluation, ethics, and GenAI concepts. The items include questions about practical use, system performance, LLMs, tokens, RAG, prompting, and responsible verification.

  • Pew Survey: The Pew AI literacy questionnaire was rigorously tested on PC and mobile devices before launch, with test data used to verify survey logic and randomization.The six-item quiz was incorporated into the paper’s broader survey.
  • AI questions: The AI items ask students to identify AI applications across customer service, email, health products, shopping, music, and other everyday contexts.These questions target recognition and practical use of AI in familiar settings.
  • Use and Apply: The survey includes applied questions about improving underperforming AI systems and identifying suitable tools for different tasks.Examples involve choosing more training data and matching tools to modalities or applications.
  • Evaluate and Ethics: The instrument also tests ethical judgment and whether students verify AI-generated information against credible references.These items address developer conduct, explainability, and the trustworthiness of LLM outputs.
  • GenAI questions: The GenAI items assess definitions, LLM operation, task capabilities, prompt-based development, tokens, and retrieval-augmented generation.The questions distinguish next-word prediction, prompt design, token processing, and supplying relevant data through RAG.

C. Chiu et al. (2024)

Chiu et al. developed and validated an objective AI-literacy test for school students rather than relying only on self-reported capability. The 25-item instrument was tested with 2,390 students and met reported reliability and validity criteria.

  • Rationale: Earlier AI-literacy studies often measured students’ perceived capability through self-report rather than actual knowledge.Chiu et al. sought an objective instrument comparable to tests used in science, mathematics, and computational literacy.
  • Instrument: The resulting school-student AI-literacy test contains 25 multiple-choice questions and was validated within a middle-school AI curriculum.Its validation involved students in grades 7 to 9.
  • Validation: The test was administered to 2,390 students and evaluated using a Rasch model.The validation results addressed dimensionality, reliability, and validity of the items.
  • Item provenance: The survey item table identifies the parent instrument for each item included in the paper’s assessment.This table documents item provenance rather than reporting student performance.

A. Course and Student Population

The survey was administered in an undergraduate Information Technology course to juniors and seniors from several computing-related areas. The student population was selected for relatively homogeneous expertise, and the course exposed students to AI through readings and case studies rather than a dedicated AI lecture.

  • Participants were undergraduate juniors and seniors enrolled in an Information Technology course.
  • Most students were minoring in cybersecurity, while others pursued networking, web development, or cloud computing concentrations.
  • The researchers sought a relatively homogeneous student population because this was the survey’s first implementation.
  • The course addressed sociotechnical and ethical technology topics, including privacy, surveillance, sustainability, and economic implications.
  • Students encountered AI through course readings and case studies, but received no dedicated AI lecture and had not previously taken any specified AI course.

B. Survey Administration

The survey was administered online as an extra-credit assignment through the institution’s examination system, with monitoring controls intended to ensure independent responses.

  • Students completed the survey online as an extra-credit assignment using the institution’s standard examination system.
  • A lockdown browser and video monitoring were used to restrict cheating or plagiarism and capture students’ own knowledge.

V. FINDINGS

Students performed well overall, but the findings revealed lower performance on application and higher-order tasks alongside misconceptions about AI mechanisms and tool-task alignment.

  • 78.83% was the average score across 30 survey items, with item performance ranging from 100% on Q6 to 33.33% on Q17.Cronbach’s alpha was 0.66, with a 95% confidence interval of [.516,.775].
  • 81.73% was the average for Know & Understand, compared with 80.99% for Ethics, 74.94% for Evaluate & Create, and 74.74% for Use & Apply.
  • The item-difficulty distribution included 8/30 items above 90%, 7/30 between 80–90%, 9/30 between 70–80%, 2/30 between 60–70%, and 4/30 below 60%.
  • 33.33% identified next-token prediction on Q17, while a majority selected the misconception that LLMs summarize the web.
  • 50.88% recognized RAG’s advantage as retrieving evidence from a corpus, while 36.84% overgeneralized it as improving generalization to unseen questions.
  • 59.65% correctly identified supplying a competitor list as the least effective applied-prompting strategy, while 22.81% undervalued structured persuasive prompting.
  • More than 80% of respondents scored above 70%, including 19.30% who scored 90% or higher.Twenty students (35.09%) scored in the 80% range, and 29.82% scored in the 70% range.
  • 11 of 30 questions were categorized as misconceptions because their correct-response rates were 75% or lower.These questions ranged from 33.33% correct on Q17 to 75.44% correct on Q43.

B. Know and Understand

Students showed strong foundational knowledge overall, but a major misconception about LLM mechanisms lowered performance on Know & Understand; application performance was moderate and exposed tool-selection difficulties.

  • B. Know and Understand: 81.73% was the mean score for Know & Understand across 12 items, indicating strong foundational knowledge.
  • B. Know and Understand: 33.33% answered Q17 correctly, revealing confusion between next-token prediction and the misconception that LLMs summarize the web.
  • B. Know and Understand: 100% answered Q6 correctly, indicating widespread mastery of that concept.
  • C. Use and Apply: 74.74% was the mean score for Use & Apply across five items, indicating partial knowledge of applying AI to practical tasks.
  • C. Use and Apply: Students handled routine and near-transfer applications better than choosing and configuring the appropriate method.

D. Evalaute and Create

Evaluate and Create performance was comparatively lower than foundational and ethical knowledge, while the survey exposed misconceptions about data, modality, and system improvement. The instrument is intended to support targeted instruction and course adaptation.

  • D. Evalaute and Create: Evaluate and Create averaged 74.94% across seven items, indicating comparatively lower performance on higher-order tasks.This category was lower than Know & Understand and Ethics in the reported category averages.
  • D. Evalaute and Create: About 70% identified a speaker as unsuitable for supplying computer-vision data, while about 20% selected a CT scanner.The response pattern suggests confusion about modality and task fit.
  • D. Evalaute and Create: Around 63% selected more training data as the best improvement lever, whereas around 30% chose increasing the learning rate.The authors characterize this as a bias toward algorithmic tweaks over data-centric improvements.
  • D. Evalaute and Create: Brief labs emphasizing modality alignment, dataset curation, and evidence-based error analysis could address these misconceptions.The proposed instructional response targets sensor relevance and primary data quality.
  • D. Evalaute and Create: The survey can identify knowledge gaps and help faculty design curricula that develop deeper comprehension of AI architectural principles.The authors describe formative and diagnostic uses, including tailoring content to deficient areas and adding domain-specific questions.

VII. LIMITATIONS AND FUTURE WORK

The study’s evidence is limited by its single-university, single-cohort setting and by the need for further validation. Future work will broaden populations, validate the instrument, and update items as AI knowledge changes.

  • VII. LIMITATIONS AND FUTURE WORK: The study used one university and one student cohort, so sampling bias may affect responses.The authors plan implementation across more disciplines, institutions, and non-student populations.
  • VII. LIMITATIONS AND FUTURE WORK: The instrument combines validated items but would benefit from additional validation studies.The authors identify future validation as a separate limitation of the curated instrument.
  • VII. LIMITATIONS AND FUTURE WORK: Because AI knowledge changes rapidly, the survey may need items added or removed to remain current.The authors explicitly recommend revisiting the instrument over time.
Loading 2608.21391v1…