Source-linked AI summary
What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development
Christopher Brooks
TL;DR
The paper asks how computational choices between AI-assisted generation and expert review shape the evidence and content available for judgment. Across two linked in-silico studies of 32,000 selected Big Five items, it follows fixed source populations through representation, structural evaluation, and candidate-form construction. Broadly stable geometry and complete forms nevertheless concealed changed item evidence, retained content, and wording, making the evaluator an inspectable part of measurement design.
Problem
AI-assisted generation can produce more plausible items than psychometricians can evaluate, creating a need to understand which computational evidence and content reach expert review.
Method
The paper varies generation packages, embedding configurations, structural methods, and candidate-form eligibility policies while tracing fixed source populations through two linked studies.
Results
Broad semantic geometry and complete quotas concealed changed construct evidence, item retention, intended content coverage, and candidate-form wording across evaluator configurations.
Takeaways & Limitations
The computational evaluator should be treated as an inspectable and revisable psychometric object, with item-level trajectories and evaluator decisions preserved for expert judgment.
Takeaways & Limitations
Evidence was derived before respondent-data validation and does not establish item functioning, reliability, respondent dimensionality, invariance, or use-specific validity.
Abstract
from arXiv · showhide
Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are usually treated as technical preliminaries. Yet representation, structural reduction, and selection policy determine which items and evidence psychometricians ever receive. Across two linked in-silico studies of 32,000 selected Big Five items, we followed fixed source populations from semantic representation through structural evaluation and candidate-form construction. Broad agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence, different items survived, and intended attributes could disappear even as community correspondence improved. These sensitivities also differed across generated source populations. At the final review boundary, both eligibility policies filled every content cell in every evaluable form, yet they presented different wording. Across embedding configurations, inclusive primary forms shared a median of only 6 of 40 items, reflecting the total downstream consequence of changing representation across structural evidence and ranking. The apparent stability of global summaries and complete forms therefore concealed instability in the content reaching psychometricians. The computational evaluator is not neutral infrastructure between generation and expertise; it is an inspectable and revisable part of measurement design.
1 Introduction
AI-assisted generation can outpace psychometric review, making computational screening consequential to which items and evidence receive attention. The studies examine how generation, representation, structural reduction, and eligibility policies shape candidate forms before expert judgment.
- Motivation: Large language models make item production inexpensive and abundant, while psychometricians cannot evaluate every plausible generated item.Generated items may be repetitive, overly evaluative, context-dependent, or aligned with neighboring constructs.
- Motivation: The computational evaluator uses embeddings, network methods, and selection rules to make large candidate pools tractable before expert review.These components also determine which content and evidence psychometricians see.
- Contribution: Candidate-form completeness does not guarantee content stability: forms can fill every position while replacing many of the items occupying those positions.Community correspondence and quota completion can likewise conceal lost intended content or changed wording.
- Study program: Study 1 varies two generators, two generation packages, five embedding configurations, and two structural methods to test effects on construct evidence, retention, and attribute coverage.Study 2 compares inclusive eligibility, retaining items supported by either method, with agreement eligibility, requiring both methods.
- Contribution: The paper argues that representation, structural reduction, and form eligibility change evidence attached to identical items, retained content, and wording reaching review.These effects can remain hidden when broad semantic geometry or form completeness appears stable.
2 Conceptual and Empirical Background
The paper situates AI-assisted item development within computational psychometrics and validity theory, emphasizing that representation and reduction choices affect what content remains available for review. It distinguishes global semantic agreement from local item retention and candidate-form consequences.
- Background: Language models extend item generation beyond templates, but inexpensive production shifts the evaluation bottleneck toward later computational and expert refinement.Prior work uses language models for construct-specific items, item relations, content balance, and refinement.
- Background: AI-GENIE provides a pathway joining generation, embeddings, structural reduction, and candidate-form construction while keeping evaluator choices inspectable.The paper separates generation eligibility, semantic representation, structural evaluation, and form eligibility.
- Semantic representation: Embedding comparisons should assess target-versus-alternative ordering across representation spaces rather than interpret raw cosine values in isolation.Similarity may reflect direct naming, diagnostic behavior, repeated definition language, or general evaluative tone.
- Semantic representation: Similar global item-pair geometry can coexist with different local neighborhoods, changing community assignments, redundancy decisions, or narrow rank cutoffs.Consequently, geometry agreement can coexist with different item retention.
- Structural evaluation: UVA and bootstrap-stable structural methods reduce item pools before final EGA, but final community structure does not establish whether intended content survived.The study compares retained-item communities with communities in the full source pool.
- Candidate forms: Inclusive eligibility preserves items retained by either method, whereas agreement eligibility requires corroboration; each policy changes the pool competing for finite form positions.Neither policy is self-justifying without examining the candidate forms produced.
3 Methods
The two-study design follows generated Big Five source populations through semantic representation and paired structural evaluation into candidate forms for psychometric review. It varies generation, embedding, structural, and eligibility choices while preserving item-level evidence and unknown outcomes.
- Study program: Study 1 covers source-population formation, semantic representation, and structural evaluation; Study 2 examines how paired evidence becomes a candidate form.Figure 1 presents the shared stages, and Table 1 summarizes each study’s role.
- Common design: The design uses 20 Big Five trait/attribute content cells as working targets for generation, anchors, coverage summaries, and candidate-form positions.The endpoint is a form for psychometrician judgment, not a finalized instrument or respondent-administered form.
- Generators and repetitions: 32,000 selected items arise from 400 generation tasks crossing two generators, two packages, five traits, and 20 independently seeded repetitions.Each task contributes an 80-item source population balanced across four attributes for one trait.
- Generators and repetitions: Generation packages differ in construct information and lexical constraints: AI-GENIE permits declared terms, whereas construct-indirect guidance prohibits them in respondent-facing items.The comparison concerns complete packages, including prompts, adaptive history, and gate rules.
- Semantic representation: Each source population is encoded separately by five embedding configurations before structural evaluation supplies method-specific evidence.The configurations jointly differ in model family, training, architecture, parameter count, dimension, and input treatment.
- Structural evaluation: The common structural sequence uses UVA, independent TMFG and graphical-lasso evaluation, bootstrap stability, and final EGA on retained items.The configured EGA procedure treats embedding-coordinate rows as cases, making vector dimension the effective sample size.
- Data and evidence handling: Generation gating removes malformed or duplicate outputs, while evidence records preserve attempted outputs, task-specific identities, trajectories, and unknown results.Unknown evidence remains explicit rather than being treated as negative evidence.
4 Study 1: Representation and Structural Evidence
Study 1 shows that representation and structural choices changed construct evidence, retained items, content coverage, and recovered structure, with effects depending on the generated source population. Broad semantic agreement therefore did not guarantee the same local evidence or surviving wording.
- Construct evidence: 35.2% of AI-GENIE source items and 25.5% of construct-indirect items changed target-first status across embedding configurations.Near-tied aggregate evidence could still attach different construct evidence to identical wording.
- Structural evidence: Qwen 0.6B, 4B, and 8B had TMFG initial AMI values of .653, .691, and .715, but only 59 of 100 construct-evidence cells were monotonic.Final AMI peaked at Qwen 4B, showing that the configured sequence was not uniformly larger-is-better.
- Source-population dependence: .080 retained-set Jaccard and .065 initial community AMI separated Qwen-generated from Gemma-generated populations in the Qwen 4B–Qwen 8B TMFG comparison.Sensitivity therefore appeared in downstream local decisions and could not be assumed to transfer unchanged across generators.
- Source-population dependence: Qwen 8B retained 46.3 percentage points more than BGE-M3 for Neuroticism/anxious content under AI-GENIE, but only 0.8 points more under construct-indirect generation.The generation package could amplify or attenuate a later representation contrast.
- Item-level illustrations: An Agreeableness/humble item switched its leading attribute across embeddings, while compassion and disciplined examples remained stable.These purposive examples show that fixed wording can acquire different construct evidence under different representations.
- Retention and coverage: .410–.565 average retained-set overlap under TMFG showed that representation changed which statements survived.Graphical lasso produced higher overlap, ranging from .666 to .755, but still retained distinct sets.
- Retention and coverage: 119 empty content cells occurred under BGE-M3 versus 27 under Qwen 8B after TMFG reduction.All tasks began with 20 items per attribute, but balanced generation did not ensure balanced survival.
- Structural evidence: .733–.874 geometry agreement contrasted with .565 retained-set overlap and .720 initial community AMI for Qwen 4B versus Qwen 8B under TMFG.Broad similarity ordering could remain stable while local retention and community structure diverged.
5 Study 2: Policy Effects and Candidate Forms
Study 2 compares inclusive and agreement eligibility policies for converting structural evidence into candidate forms. Both policies completed balanced forms, but agreement removed substantial eligibility and changed the exact wording presented for expert review.
- 5.1 Purpose and Research Question: The inclusive policy accepted items retained by either structural method, whereas agreement required retention by both.The comparison isolates form eligibility as a policy applied after shared structural evaluation.
- 5.2 Candidate-Form Construction: Items were ranked by combined target advantage for attribute definitions and behavioral indicators before joint assignment to 20 content cells.For the construct-indirect package, these anchors also shaped generation, so the score was package-aligned rather than independent quality evidence.
- 5.3 Results: 53.1% of complete paired outcomes were retained by both methods, while 25.0% were graphical-lasso-only and 3.8% TMFG-only.The two policies differed only on one-method outcomes, which comprised 28.8% of complete pairs.
- 5.3 Results: 46,121 of 131,073 complete item occurrences were removed by agreement, contracting the inclusive pool by 35.2%.Across complete configurations, agreement removed roughly one fifth to more than one half of the inclusive pool.
- 5.3 Results: Both policies filled all 40 primary and 40 alternate positions in every evaluable form, preserving all 20 content cells.Completeness therefore did not imply identical wording or identical candidate evidence.
- 5.3 Results: Matched primary forms shared a median of 29 of 40 items, leaving a median 11 items unique to each policy.Including alternates, a median 21 items in each policy’s form were absent from the other.
- 5.3 Results: The policies showed no consistent advantage in warning burden, local-dependence evidence, or source-task concentration.The demonstrated consequence was changed wording presented for review rather than uniformly improved diagnostics.
- 5.3 Results: Inclusive forms resembled graphical-lasso-only forms, whereas agreement forms resembled TMFG-only forms because TMFG was more selective.This explains realized form differences without establishing that either structural method produced better items.
6 General Discussion
The evaluator’s representation, structural, and form-construction choices alter the construct evidence, content, and wording that reach psychometrician review. Broad stability and complete forms can therefore conceal substantial dependence in the exact review material.
- Evaluator dependence: The evaluator is not neutral: computational choices determine which items and evidence survive to expert review.The studies held source populations fixed for downstream contrasts and traced differences through candidate-form assembly.
- Property-specific stability: Much overall semantic ordering remained stable, but exact retained-item overlap and intended content coverage were less stable.Community correspondence could improve while intended attributes disappeared from the surviving material.
- Representation: Construct evidence changed with both embedding representation and the anchor family used to define target comparisons.The comparison set is therefore part of construct representation, not merely a scoring aid.
- Structural and content evidence: Structural methods often strengthened correspondence with intended attributes while reducing content coverage, so structural and content evidence remained connected but non-interchangeable.Reviewers need both structural trajectories and content information to distinguish dependence-related removal from construct concerns.
- Candidate forms: Both eligibility policies produced complete, balanced forms in all 19 evaluable configurations, yet they presented different items for judgment.Global rank agreement and eligible-pool overlap exceeded overlap among the few items occupying each content cell.
- Traceability: Traceability preserves stage histories and displaced candidates, allowing the evaluator to be compared and revised without stochastic regeneration confounding later contrasts.Final item lists alone cannot reveal whether content was absent, structurally removed, policy-excluded, or displaced at the form boundary.
- Scope: These studies establish dependence within the configured evaluator, not the quality or validity of a finished instrument.Respondent data, psychometrician review, and validation studies remain necessary.
7 Limitations and Future Directions
The evidence is bounded by the evaluated constructs, models, configurations, and pre-validation design. Future work must test other settings, isolate consequential components, and examine outcomes with psychometricians and respondents.
- Scope boundaries: The findings come from one Big Five framework, two generators, five embedding configurations, two structural methods, two eligibility policies, and one form procedure.Other constructs, languages, formats, models, anchor systems, pool sizes, methods, or review constraints may yield different patterns.
- Configuration confounding: Generator and embedding comparisons do not isolate individual model or representation properties because their configurations differ in bundled characteristics.The results should therefore be interpreted as dependence on complete configurations rather than general model rankings.
- Validation boundary: Pre-validation evaluator evidence does not establish item functioning, reliability, respondent dimensional structure, measurement invariance, or validity.These claims require psychometrician review, response-process evidence, and appropriately designed respondent studies.
- Uncertainty and noncompletion: Resampling captures variation across repeated generation tasks, not uncertainty from participant sampling.One structurally incomplete configuration remained non-evaluable where paired evidence was required, preserving absent evidence as distinct from negative evidence.
- Form capacity: Complete forms depended on large source populations and four items per content cell, so smaller pools or stricter constraints may underfill cells.The calibration supports a pragmatic workflow choice, not a general optimum.
8 Open Science, Data, and Materials
The analysis-ready data and code are publicly available through Zenodo, together with the materials needed to inspect generation, evaluation, outputs, and provenance.
- Open materials: Analysis-ready data and Python code are publicly available through Zenodo at 10.5281/zenodo.21968239.The deposit includes prompts, model responses, selection records, embeddings, structural evidence, form outputs, provenance documentation, and pinned software requirements.
9 Conclusion
Across two studies, source-population formation and evaluator configuration shaped the evidence, items, and content reaching psychometricians. The evaluator should therefore be treated as an inspectable part of measurement design rather than neutral infrastructure.
- 9 Conclusion: Both candidate-form policies filled every content cell in all 19 evaluable configurations, but they presented different items for judgment.Changing embedding configuration produced greater observed candidate-form divergence than changing form policy.
- 9 Conclusion: Representation changed the construct evidence attached to fixed items and the content surviving structural reduction.These effects linked source-population formation, semantic representation, structural evaluation, and candidate-form construction.
- 9 Conclusion: The evaluator determines what screening represents, removes, includes in candidate forms, and makes inspectable before expert review.Preserving item identity, semantic contrasts, structural trajectories, policy decisions, displaced content, and unknown evidence supports studying and improving the evaluator.
- Construct framework: The construct framework organized the evaluator around five traits and 20 working attributes rather than proposing a facet taxonomy.Definitions, behavioral indicators, boundaries, and confounds delimited the intended content domain.
Appendix B. Evaluator Configuration and Candidate-Form Construction
The appendix defines the evaluator’s configured representations, outcome units, structural traces, eligibility policies, and resampling boundaries. These choices preserve distinctions among semantic evidence, structural states, coverage, and candidate-form content.
- B.1 Configured Model Panel: The evaluator comprises bundled embedding configurations whose model properties cannot be isolated into single component effects.Qwen comparisons provide bounded within-family and same-dimension diagnostics, not component-causal tests.
- B.2 Outcome Definitions and Units: Outcome definitions use local denominators across item-, task-, structural-evaluation-, and form-level units.Geometry, identity, community correspondence, and coverage remain distinct rather than forming one general item-quality measure.
- B.3 Structural Evaluation: TMFG and graphical lasso were applied separately after a shared UVA step to produce method-specific structural evidence.Embedding-coordinate rows served as EGA cases, making vector dimension the effective sample size for regularization and stability procedures.
- B.3 Structural Evaluation: Structural traces distinguish UVA removal, stability-threshold removal, retention, and unknown evidence.Missing or incomplete evidence remains distinct from removal, and final stability does not retroactively change form-construction decisions.
- B.4 Paired Method States and Eligibility Policies: The inclusive policy accepts retention by either method, whereas the agreement policy requires retention by both methods.Both require complete paired evidence and use the same anchor ranking, uniqueness constraint, assignment, and quotas.
- B.4 Paired Method States and Eligibility Policies: Secondary method-specific forms were diagnostic and did not define additional primary policies or establish a preferred structural method.The focal policies were rebuilt from the common draw across embedding configurations.
- B.7 Generation-Task Resampling: Task resampling preserves provenance and form-level outcomes but reflects dependence on repeated generation-task mixtures, not participant sampling.Comparisons involving missing structural evidence remain nonevaluable.
- Evaluator consequences: Across representations, identical wording could receive different target-versus-competitor evidence and different structural outcomes.Target-first status varied for 35.2% of AI-GENIE-source items and 25.5% of construct-indirect items across embeddings.
C.2.1 Selected Task-Level Comparisons
The task-level comparisons show that source population and generation package moderated how representations preserved items, communities, and anchor alignment. These contrasts localize dependence within the configured design without ranking generators or establishing item validity.
- C.2.1 Selected Task-Level Comparisons: These comparisons localize dependence within the configured design rather than ranking generator quality, isolating prompt effects, or establishing item validity.The unit is the generation task or paired source population named in each contrast.
- C.2.1 Selected Task-Level Comparisons: Source population moderated how consistently representations preserved items and communities across task-level comparisons.The contrasts used planned, observed, and pair-complete generation tasks with shared bootstrap draws.
- C.2.1 Selected Task-Level Comparisons: Compound generation packages changed the kind of anchor alignment visible in generated content.Because instructions and candidate gating changed together, the contrast cannot be reduced to label prohibition or behavioral guidance alone.
C.3 What Structural Reduction Preserved and Lost
Structural reduction generally increased correspondence with intended attributes while changing retained item populations and sometimes losing content coverage. Complete candidate forms therefore did not guarantee stable wording, full attribute representation, or complete evidence.
- C.3 What Structural Reduction Preserved and Lost: Exact four-community solutions rose from 32.93% initially to 73.02% finally after reduction.AMI decreased in 59 of 1,999 evaluable TMFG contexts and 217 of 2,000 graphical-lasso contexts.
- C.3 What Structural Reduction Preserved and Lost: Under TMFG, BGE-M3 had 119 empty attribute cells among 1,596 evaluable cells, versus 27 of 1,600 under Qwen 8B.Graphical-lasso zero coverage ranged from 0.2% under Qwen 4B to 1.3% under BGE-M3.
- C.3 What Structural Reduction Preserved and Lost: AMI, ARI, and NMI agreed on all 45 aggregate pairwise orderings for change, but differed on one initial and two final orderings.At the task level, they assigned the same sign to change in 3,882 of 3,999 evaluable contexts.
- C.3 What Structural Reduction Preserved and Lost: One BGE-M3/TMFG task had incomplete structural evidence, leaving its context and 80 item outcomes unknown rather than rejected.The missing result was one analytic noncompletion, not 80 rejected items.
- Candidate-form construction: Agreement removed 42,830 of 124,408 inclusive-eligible item occurrences, a pooled reduction of 34.4%.Configuration-specific reductions ranged from 21.4% to 53.0%, yet both policies filled every primary and alternate position in every evaluable configuration.
- Candidate-form construction: Complete content-cell coverage coexisted with different exact wording across the two eligibility policies.The forms describe what the configured evaluator placed before psychometricians, not which policy produced higher-quality items.
D.2 What Persisted Across Generation-Task Mixtures
Across repeated generation-task mixtures, evaluator outputs remained complete but wording and sensitivity varied substantially with embedding configuration, eligibility policy, and source-population depth. The resampling analysis also identified a configuration-specific non-evaluability boundary.
- Resampling evidence: 19,359 of 20,000 planned paired configuration draws were evaluable, and every evaluable draw filled all 40 primary and 40 alternate positions under both policies.Configuration-specific median primary-wording Jaccard ranged from .379 to .778.
- Wording stability: 36 of 40 evaluable embedding pairs shared a median of 6 of 40 primary statements within the same source population and eligibility policy.Their median Jaccard was .081, compared with 29 of 40 shared statements and a median Jaccard of .569 for the 19 evaluable inclusive-versus-agreement comparisons within configurations.
- Policy diagnostics: Inclusive and agreement forms averaged 16.53 and 15.79 warning-bearing items, respectively, while primary forms drew from approximately 34 of 100 source tasks under either policy.No single task supplied more than three of 40 primary statements.
- Source-population depth: Source-population depth can alter redundancy decisions, community assignments, retention, and rankings because adding items changes the population against which candidates are evaluated.The study therefore treated depth as a measurement-design choice before fixing the 80-item source populations.
- Source-population depth: The 80-item populations were formed from ordered, gate-eligible sets of 100 items, balanced at 25 items per attribute, using exact prefixes of 40, 60, 80, and 100 items.The ordering recorded eligibility timing rather than estimated quality rank.
- Evaluation design: Each depth was evaluated separately across five embedding spaces and two structural methods, with adjacent comparisons restricted to complete final-item evidence.Paired contexts numbered 950 for 40-60, 959 for 60-80, and 958 for 80-100; graphical lasso completed all 480 contexts at every depth.
E.2 Evidence Across Adjacent Depths
Adjacent source-population depths produced strong agreement in communities for shared retained items, but narrow diagnostic slates remained sensitive. The combined descriptive evidence supported 80 items as a configured operating point rather than a universal optimum.
- E.2 Evidence Across Adjacent Depths: Retained populations became more similar across successive depth contrasts, while the mean number of newly retained items declined from 15.42 to 10.86.Mean retained-item Jaccard changed differently by method between 60-80 and 80-100: .4670 to .4526 for TMFG and .5989 to .6575 for graphical lasso.
- E.2 Evidence Across Adjacent Depths: Community assignments for shared retained items agreed closely throughout, but broader retained-set similarity did not imply stability at the narrow diagnostic-slate boundary.The diagnostic slates remained sensitive at every depth.
- E.2 Evidence Across Adjacent Depths: Approximately one third of larger-pool-only slate entries came directly from bands added in the 40-60 and 60-80 expansions.In the 80-100 expansion, direct contributions from new ranks fell, while later additions became more similar to earlier candidates.
- E.3 Why 80 Items Was Selected: The combined evidence supported 80 items, balanced as 20 gate-eligible statements per attribute, as a pragmatic source-population depth for the configured studies.Stopping at 60 would exclude ranks 16-20, whereas expanding from 80 to 100 continued altering the evaluator but supplied a smaller share of new slate content.
- E.3 Why 80 Items Was Selected: The depth decision was descriptive and did not fit a formal elbow estimator, test equivalence, or apply a universal similarity or community-agreement threshold.It estimated sensitivity across exact eligibility-order prefixes from two traits, two generators, two generation packages, five embedding spaces, and two structural methods.
- E.3 Why 80 Items Was Selected: Transferring 80 items to all five traits and the primary studies’ 20-repetition design was a configured workflow choice, not a separately estimated generalization.The results do not establish that 80 items are universally optimal, exhaust a trait’s content, or produce respondent-valid scales.