Source-linked AI summary
GenAIT: Development and Validation of an Objective Generative AI Literacy Test for High School Students
Brett Puppart, Kristjan-Julius Laak, Jaan Aru
TL;DR
Objective GenAI literacy assessment among high school students remains underdeveloped. This study developed and validated an Estonian-language objective GenAI literacy test, finding evidence supporting its validity while indicating greater suitability for group-level research than individual classification.
Problem
Objective GenAI literacy assessment among high school students remains underdeveloped, motivating development of an Estonian-language test.
Method
The study developed GenAIT through multiple stages informed by educational and psychological testing guidelines.
Results
Validity evidence supported GenAIT’s test content, internal structure, and relations to external variables.
Takeaways & Limitations
GenAIT provides an objective measure of GenAI literacy for group-level research among high school students.
Takeaways & Limitations
GenAIT scores are more suitable for group-level research than for high-stakes individual classification.
Abstract
from arXiv · showhide
There is growing international interest in generative AI (GenAI) literacy and its assessment among high school students, but objective assessment in this population remains underdeveloped. This article reports the iterative development and validation of the GenAI Literacy Test (GenAIT), an 18-item multiple-choice test measuring high school students' conceptual knowledge about GenAI, with content spanning technical, practical, and human-impact domains. Expert review of relevance, clarity, and comprehensiveness provided evidence of content validity. In a large-scale survey of 7432 Estonian high school students, we evaluated the psychometric functioning of the Estonian-language GenAIT using confirmatory factor analysis, classical test theory, and item response theory. Results supported approximate unidimensionality, broadly adequate reliability for group-level research (marginal reliability = .72, KR-20 = .69), and good fit of a three-parameter logistic model (RMSEA = .013, TLI = .987, CFI = .990, SRMSR = .021). Measurement precision was sufficient for the majority of students but varied substantially across the latent trait, with lower precision for lower scoring students. GenAIT is therefore more suitable for group-level research than high-stakes individual classification. GenAIT scores were unrelated to perceived usefulness and perceived ease of use, and negatively associated with LLM use frequency, suggesting that frequent use and favorable perceptions of AI should not be treated as proxies for conceptual understanding.
1. Introduction
GenAI is reshaping how people live, work, and learn, creating a need for AI education and objective measures of students’ critical and ethical understanding. Existing assessments inadequately target GenAI, often rely on self-report, and lack validation for high school students; this study addresses these gaps with an objective Estonian-language test.
- Assessment need: Objective measurement is needed to assess students’ real GenAI knowledge and understanding, rather than relying mainly on subjective self-reports.Self-report assessment may produce inaccurate estimates of actual knowledge and understanding.
- Study aim: The study addresses these methodological gaps by developing and validating an objective GenAI literacy test in Estonian.The test is intended for the high school population, which is also targeted by large-scale initiatives such as the Estonian AI Leap.
2. Background
GenAI literacy is framed as a focused extension of AI literacy, emphasizing conceptual knowledge across technical, practical, and human-impact domains. Existing assessments remain limited for K–12 students, motivating development of an objective Estonian-language test for high school students.
- Conceptualization: GenAI literacy extends AI literacy by addressing generative technologies’ distinctive reliance on natural-language prompting and risks such as over-reliance and misinformation.The construct is treated as focused rather than wholly distinct from AI literacy.
- Conceptualization: The GenAIT framework spans technological understanding, effective use, and human impacts on cognition, society, human rights, democracy, the rule of law, and the environment.These domains were used to ensure breadth rather than to define three separate dimensions, because they were expected to be intertwined.
- Conceptualization: The technological domain covers how LLMs work, the practical domain covers prompting and critical evaluation, and the human domain covers cognitive, societal, and environmental implications.Examples include token prediction and model scaling, prompt quality and hallucinations, and risks to learning, democracy, information ecosystems, inequality, and environmental sustainability.
- Assessment gap: Validated AI-literacy instruments exist, but generative-AI knowledge assessments remain scarce and especially limited for K–12 students, with most relying on self-report questionnaires.This gap is particularly relevant to Estonia’s national AI Leap initiative and the need for objective assessment.
- Study contribution: GenAIT was developed as an objective Estonian-language measure of high school students’ conceptual GenAI knowledge, with 18 multiple-choice items covering technical, practical, and human domains.The instrument is primarily intended for group-level research, using both summative raw scores and IRT-derived latent trait scores.
3. Method
GenAIT was developed and validated through staged procedures aligned with educational and psychological testing guidelines, combining expert-panel content evidence, pilot external-variable evidence, and research-project analyses of internal structure and reliability. Its content domain was defined as technical, practical, and human GenAI literacy, with concepts selected for age appropriateness and lasting educational relevance.
- 3. Method: GenAIT development and validation used staged evidence from expert review, a pilot study, and research-project analyses of internal structure and score reliability.The procedures followed educational and psychological testing guidelines; the research-project data came from evaluating the AI Leap initiative.
- 3. Method: The development process documented item retention across stages and presented domain concepts and sample items in Table 1.The process overview was provided in Fig. 2, while Table 1 introduced the content domains, concepts, and sample items.
- 3. Method: GenAI literacy was operationalized across technical, practical, and human domains to ensure representative content coverage, not separate latent dimensions or subscale scores.The domains were specified as content categories rather than as an assumption about the test’s latent structure.
- 3. Method: The initial content pool comprised 21 target concepts, selected through literature review and author discussions for relevance, age appropriateness, and durable educational value.Selection prioritized concepts expected to remain relevant despite rapid technological change rather than current model-specific capabilities.
C. It cannot create new paintings
This section presents Item 14, which asks why large language models produce overly general and neutral essays. The item concerns how prompting relates to model outputs and capabilities.
- Item 14: Item 14 asks what a person is most likely doing wrong when an LLM produces overly general, neutral essays.The supplied passage identifies the scenario but does not provide the correct answer option.
- Item 14: The surrounding concepts include prompt-induced bias and prompt-dependent capabilities.
B. They are writing prompts that are too vague
GenAIT was developed through iterative item construction, expert review, and pilot testing, yielding evidence of strong content validity, adequate preliminary reliability, and expected associations with magical perceptions of AI. The final analytic sample comprised 7,432 respondents.
- B. They are writing prompts that are too vague: Items were iteratively drafted, reviewed, and refined by five Artificial and Natural Intelligence Lab researchers at the University of Tartu.Each Estonian-language item used four answer options with one correct response, and disagreements were resolved through discussion and consensus.
- B. They are writing prompts that are too vague: Most items assessed conceptual understanding through realistic GenAI scenarios rather than isolated factual recall, with distractors targeting misconceptions or partially correct intuitions.The item-generation philosophy resembled concept inventories, but common misconceptions were not systematically identified and validated (Furrow & Hsu, 2019).
- 3.2. Expert review: Expert review used relevance, clarity, and comprehensiveness ratings from five AI, data science, and educational technology experts.Experts evaluated Estonian-language items using four-point forced-choice scales and provided qualitative feedback on missing concepts and item quality.
- 3.3. Pilot study: Pilot reliability was adequate (KR-20 = .77, 95% CI [.69, .84]), but Item 15 showed poor discrimination and a low item-total correlation.Distractor analysis identified a dysfunctional distractor, which was revised before the main study.
- 3.3. Pilot study: −.30 was the final 18-item GenAIT’s correlation with magical perception of AI (95% CI [−.50, −.07], p = .010, n = 71), supporting preliminary validity.The 20-item version showed a stronger negative association, Spearman’s ρ = −.35, 95% CI [−.54, −.13], p = .003, n = 71.
- 3.4.4. Data preparation and statistical analyses: 7,432 respondents remained in the final analytic sample after exclusions based on listwise and engagement-related criteria.Engagement exclusions included a response-time criterion.
4. Results
Results supported approximate unidimensionality and a three-parameter logistic model for the 20-item GenAIT, while item analysis led to removing Items 3 and 10. Most remaining monotonicity deviations were minor and concentrated at score extremes.
- Model fit: The 20-item GenAIT showed approximate unidimensionality, and the 3PL model fit best among candidate IRT models with local independence supported.The one-factor CFA yielded RMSEA = .023, TLI = .957, CFI = .962, and SRMR = .045; maximum Yen’s Q3 = .045.
- Item analysis: Items 3 and 10 were removed after item analysis identified psychometric weaknesses and conceptual redundancy.Item 3 had negative CFA loading and IRT discrimination, a significant S-X2 statistic, and lacked monotonicity; Item 10 had weak CFA loading despite strong IRT discrimination and overlapped with prompting-related items.
- Item analysis: Most observed deviations from strict monotonicity were minor and occurred at score extremes where few respondents were represented.Other flagged items were retained after content-validity considerations.
Appendix E.
The 18-item GenAIT showed acceptable 3PL fit, approximate unidimensionality, and local independence across development and validation subsamples. Reliability and precision were adequate for group-level measurement but weaker at lower trait levels, while scores were negatively associated with LLM-use frequency and unrelated to perceived usefulness or ease of use.
- Validation model fit: The validation analysis supported approximate unidimensionality and no problematic local dependence, with RMSEA = .025, TLI = .958, CFI = .963, SRMR = .047, and maximum Q3 = .035.Some items showed minor monotonicity deviations, mainly at trait extremes with few respondents.
- Validation model fit: RMSEA = .013, TLI = .987, CFI = .990, and SRMSR = .021 indicated acceptable 3PL fit in the validation subsample, despite a significant M2 test.The M2 test was M2 = 193.267 (117), p < .001.
- Item characteristics: Discrimination parameters ranged from 0.59 to 4.23, item difficulties from −1.71 to 2.05, and pseudo-guessing parameters from 0.00 to 0.33.The item set therefore covered a broad range of the latent trait.
- External associations: GenAIT scores correlated negatively with LLM-use frequency, while adjusted correlations with perceived usefulness and perceived ease of use were nonsignificant.Raw scores and latent-trait estimates were very strongly associated, Spearman’s ρ = .97, 95% CI [.97, .97], p < .001.
1. GenAIT 8.35 3.39 —
The reported results table lists perceived ease of use with a mean of 13.43, a standard deviation of 3.00, and correlations of .01, .36***, .33***, and .62***.
- Perceived ease of use: 13.43 was the reported mean for perceived ease of use, with a standard deviation of 3.00.The table also reports correlations of .01, .36***, .33***, and .62***, although the corresponding variables are not identified in the supplied passage.
5. Discussion
The discussion presents GenAIT as a promising objective measure of high school students’ conceptual GenAI knowledge for exploratory group-level research and educational evaluation. Its validity evidence is encouraging but qualified by incomplete construct coverage, uneven measurement precision, and weak correspondence between objective literacy, perceptions, and use frequency.
- Content validity: Expert feedback supported item relevance and clarity, but suggested that statistical modeling and ethics were underrepresented, so GenAIT is not an exhaustive measure of GenAI literacy.The current form captures important themes, while future versions should broaden construct coverage and add easier items for lower-literacy students.
- Psychometric performance: GenAIT showed approximate unidimensionality, acceptable item functioning, and broadly adequate fit, supporting a single score summarizing technical, practical, and human GenAI knowledge.However, measurement precision was higher for average and above-average students than for lower-scoring students, limiting high-stakes individual interpretation.
- Reliability and precision: GenAIT reliability was sufficient or near sufficient for group-level research, but its information was concentrated at average and higher ability levels rather than across the observed distribution.This mismatch makes the test better suited to group comparisons and educational evaluation than high-stakes individual interpretations.
- External variables: Objective GenAI literacy was negatively associated with LLM use frequency but unrelated to perceived usefulness and ease of use, indicating that subjective perceptions and adoption are not proxies for conceptual understanding.Perceived usefulness and ease of use were positively associated with LLM use frequency, whereas objective literacy showed a different pattern.
- Educational implications: Students struggled particularly with technical items despite frequent LLM use, supporting instruction on foundational mechanisms, verification, limitations, and ethical and societal considerations.Educational use should preserve students’ agency, cognitive engagement, and self-regulation rather than treating tool operation or prompting as sufficient literacy.
6. Limitations
GenAIT has important limitations in measurement precision, validity evidence, subgroup generalizability, and coverage of the evolving GenAI-literacy construct. Future work should improve lower-trait precision, evaluate performance-based and subgroup validity, examine response processes, and update content through repeated expert review.
- Measurement scope and precision: Lower precision below the mean and focus on conceptual knowledge make GenAIT better suited to exploratory group comparisons than high-stakes individual classification.GenAIT does not show whether students can apply knowledge critically and effectively in authentic GenAI-assisted tasks.
- Validity evidence: GenAIT needs easier items for lower-trait precision and evaluation against performance-based measures such as prompting and output evaluation.Evidence from external-variable associations is limited because the pilot was small and potentially selective, while the main study lacked objective and performance-based criteria.
- Subgroup generalizability: Missing demographic data prevented detailed sample description and tests of differential item functioning or measurement invariance across grade level, gender, and home language.Despite adequate overall IRT fit and no problematic local dependence, generalization across student subgroups remains unestablished.
- Content and response processes: Item misconceptions were based on literature and researcher judgment without directly examining student interpretations, while content necessarily covered only part of an evolving construct.Cognitive interviews, think-aloud protocols, renewed expert review, and periodic item revision could improve response-process evidence and content coverage.
7. Conclusion
GenAIT is an 18-item objective test of high school students’ conceptual GenAI knowledge with initial validity and reliability evidence. Its precision was higher for average and above-average literacy, while usage frequency and favorable AI perceptions did not indicate conceptual understanding.
- 7. Conclusion: GenAIT is a promising foundation for assessing high school students’ conceptual GenAI literacy, supported by validity evidence from content, internal structure, external relations, and score reliability.The study developed the 18-item objective test and provided initial validity evidence.
- 7. Conclusion: GenAIT measured average and above-average GenAI literacy more precisely than lower levels, indicating a need for easier items in future versions.Measurement precision varied across the literacy continuum, with lower precision at lower levels.
- 7. Conclusion: Objectively measured GenAI literacy was not positively associated with perceived usefulness or ease of use and was negatively associated with frequency of LLM use.These findings distinguish conceptual understanding from AI adoption and favorable perceptions of AI.
- 7. Conclusion: GenAIT therefore distinguishes AI adoption and favorable AI perceptions from demonstrated conceptual understanding.Favorable perceptions and frequent use should not be treated as indicators of conceptual GenAI literacy.
- 7. Conclusion: Further validation, refinement, and cross-cultural replication are needed to strengthen GenAIT’s foundation for GenAI literacy assessment.The conclusion presents GenAIT as promising but not yet definitive across contexts.
Declaration of generative AI use
The authors used ChatGPT during manuscript preparation to discuss ideas, clarify concepts, support methodological and theoretical understanding, assist with translation and language improvement, and enhance readability. They critically reviewed and edited the resulting content and take full responsibility for the published article.
- Declaration of generative AI use: ChatGPT supported idea discussion, concept clarification, and understanding of methodological and theoretical topics.
- Declaration of generative AI use: The authors also used ChatGPT for translation and improving the manuscript’s language and readability.
- Declaration of generative AI use: After using ChatGPT, the authors critically reviewed and edited the content and accepted full responsibility for the published article.