Source-linked AI summary
AI for Computational Design Science: A Responsible Human-AI Framework and Case Study on Short-Form Video Safety Surveillance
Wenli Zhang, Jiaheng Xie, Zhihe Pan, Yidong Chai, Xiao Fang, Sudha Ram
TL;DR
Existing computational design science literature does not specify how AI participation should be organized across the CDS lifecycle. The paper introduces AI4CDS and instantiates it through ChildRiskGuard, whose architecture satisfies four stated axioms.
Problem
Existing CDS literature does not specify how AI participation should be organized across the CDS lifecycle when AI becomes a constitutive participant in producing artifacts and design knowledge.
Method
The paper develops AI4CDS as a framework for CDS and instantiates it through ChildRiskGuard, including knowledge-anchored mechanism structuring and residual learning beyond a generic-safety channel.
Results
ChildRiskGuard’s proposed architecture satisfies four stated axioms, while its exponential link is identified as the unique continuous strictly increasing link up to positive rescaling.
Takeaways & Limitations
AI4CDS is presented as the primary framework contribution, with ChildRiskGuard serving as its computational design-science instantiation.
Takeaways & Limitations
Existing CDS literature does not specify how AI participation should be organized across the CDS lifecycle.
Abstract
from arXiv · showhide
Artificial intelligence (AI) is transforming not only what information systems researchers design, but also how design research is conducted. Yet existing literature offers limited guidance for computational design science (CDS) when AI actively participates in problem formulation, resource construction, design search, evaluation, and knowledge abstraction. We develop AI for Computational Design Science (AI4CDS), a five-phase methodological framework in which AI expands problem and design search while researchers retain responsibility for domain grounding, admissibility, verification, and scientific judgment. Collaboration is governed by graduated trust, reversibility, auditability, and differentiated reproducibility. We instantiate AI4CDS through ChildRiskGuard, an interpretable artifact for detecting short-form videos inappropriate for children, while documenting AI interactions, rejected alternatives, corrections, and audit trails. The case translates audience-dependent safety and explanation faithfulness into three technical challenges and develops an artifact that separates generic from child-specific risk, represents distinct developmental-risk mechanisms, and makes concept-level explanations part of the predictive computation. ChildRiskGuard achieves an F1 score of 0.769, substantially outperforming direct application of a general-purpose content-safety model while remaining competitive with strong benchmarks. The primary contribution is AI4CDS as a responsible framework for AI-enabled CDS; ChildRiskGuard provides process and artifact evidence of how AI-expanded, researcher-governed design can generate and evaluate novel computational design knowledge.
1. Introduction
AI4CDS proposes a five-phase framework for governing AI participation across computational design science while retaining human responsibility for domain grounding, admissibility, verification, and scientific accountability. ChildRiskGuard instantiates the framework by translating child-oriented safety requirements into an interpretable short-video surveillance artifact.
- Research gap: Existing CDS guidance does not specify how AI participation should be organized across the research lifecycle.The gap concerns AI involvement in problem formulation, resource construction, design search, evaluation, and knowledge abstraction.
- AI4CDS framework: AI4CDS structures responsible human–AI collaboration through five phases spanning problem formulation, resource construction, design search, evaluation, and design-principle distillation.AI expands scientific search, while researchers determine admissibility, verify evidence, and make scientific commitments.
- AI4CDS framework: Graduated trust, reversibility, auditability, and differentiated reproducibility govern collaboration across the AI4CDS lifecycle.These mechanisms constrain how AI-generated alternatives and evidence are used across research activities.
- ChildRiskGuard case: Child-oriented short-video safety requires audience-sensitive design because general-audience acceptability does not guarantee developmental appropriateness.General-purpose harmful-content detection therefore cannot be directly transferred without translating domain knowledge into computational requirements.
- ChildRiskGuard case: ChildRiskGuard addresses audience-dependent safety and explanation gaps by preserving generic safety, isolating child-specific risk, representing developmental mechanisms, and ensuring concept-level faithfulness.The resulting artifact is an interpretable safety surveillance system for short-form videos.
- Contributions: The study contributes AI4CDS as its primary framework and uses ChildRiskGuard as consequential artifact and process evidence for AI-enabled computational design science.The case documents how domain knowledge becomes computational requirements that guide interpretable system design.
2. Related Work
Prior work provides useful safety detection and explanation techniques, but does not adequately represent audience-dependent child risk or ensure that explanations reflect the computation producing predictions. The paper therefore motivates a concept-based design that combines generic-safety references with externally grounded developmental-risk mechanisms.
- Audience-dependent safety: Existing moderation systems provide generic-safety information but cannot serve as proxies for child appropriateness.Child-oriented detection must account for risks arising specifically because the audience is children, rather than merely lowering a general detector’s threshold.
- Audience-dependent safety: Existing detection approaches generally do not distinguish generic safety risk from risk arising specifically for children.These approaches define risk from content without explicitly representing the audience for whom the content is harmful.
- Faithful explanation: Vision-language rationales may not reflect the evidence that actually drove a moderation decision because rationale generation is separate from prediction.This motivates explanations derived from the same quantities that produce the decision.
- Concept-based prediction: Concept-based methods place named concepts on the predictive pathway so prediction and explanation operate over shared intermediate variables.The paper adopts this foundation because concepts can represent developmental mechanisms directly and participate in the computation producing the decision.
- Faithful explanation: Concept bottlenecks require control over both information entering through a named concept and the concept’s contribution to the target prediction.Merely including human-interpretable concepts does not guarantee faithful explanations because concepts may encode unintended information.
- ChildRiskGuard design: ChildRiskGuard uses vision-language representations for semantic measurements while named developmental concepts structure child-specific risk assessment.The architecture separates generic-safety reference information from additional developmental-risk mechanisms and supports an extensible set of externally grounded mechanisms.
3. Research Design
The research design combines AI-augmented resource construction and design search with researcher-led validation, selection, and documentation. ChildRiskGuard formulates child-oriented moderation as a two-stage, interpretable risk-assessment problem that separates generic safety from child-specific mechanisms.
- AI-Augmented Data & Resource Construction and Design Search: AI explores resources and generates component alternatives, while researchers test validity, inspect measurements, assign signals, and freeze training-data normalization and thresholds.Candidates are paired with ablations, diagnostics, or theoretical tests before researchers retain, revise, or reject them.
- Computational Problem Formulation: ChildRiskGuard addresses child-oriented moderation as both prediction and explanation, because child-inappropriateness overlaps with but differs from generic content safety.Some risks are shared with generic safety, whereas others arise from child-specific developmental mechanisms.
- Computational Problem Formulation: The architecture separates a frozen measurement stage from a trainable risk-assessment stage that transforms generic-safety readings, concept measurements, and evidence channels into child-oriented risk.Only the risk-assessment function is learned from child-inappropriateness labels; the measurement mapping remains fixed during training.
- ChildRiskGuard Architecture: ChildRiskGuard uses externally grounded mechanisms and bounded, zero-anchored, nondecreasing response functions to preserve interpretable and directionally constrained risk contributions.Piecewise responses capture nonlinear effects such as thresholds and saturation, while the architecture supports additional externally grounded mechanisms.
- ChildRiskGuard Architecture: The retained design coordinates generic-risk separation, developmental-risk mechanism structuring, semantic-budget constraints, and concept-level additive attribution.These elements were selected after alternatives were evaluated against admissibility criteria and are integrated into a single risk-assessment function.
- Identifiable Child-Specific Residual Learning: Up to positive rescaling, the exponential link is the unique continuous strictly increasing link satisfying additive decomposition and noisy-OR composition.This hazard representation preserves OR semantics on the probability scale while keeping risk sources additive and nonnegative on the hazard scale.
Appendix B.3.3.
The appendix specifies how semantic budgets, monotone mappings, joint optimization, and hazard decomposition implement faithful concept-level prediction. It also establishes structural faithfulness properties while limiting exact Shapley equivalence to the additive regime.
- Semantic Contribution Budget: The semantic contribution budget applies to the mechanism-intensity pathway, while contextual attenuation concepts operate only on the generic-safety channel.Activation-family effects are governed separately by the knowledge-anchored gate.
- Axiomatically Faithful Monotone Additive Hazard Mapping: The monotone additive mapping keeps realized concept contributions observable while allowing nonlinear responses such as thresholds and saturation.A concept’s contribution can remain directly represented rather than hidden inside an unconstrained nonlinear mapping.
- Axiomatically Faithful Monotone Additive Hazard Mapping: Under the stated hazard, mechanism, and mapping assumptions, the child-specific pathway satisfies four faithfulness requirements as structural properties of the artifact.The formal definitions and complete verification are provided in Online Appendix B.3.4.
- Axiomatically Faithful Monotone Additive Hazard Mapping: Conditional on a realized mechanism gate, each concept contribution coincides with its Shapley value in the corresponding additive intensity game.When semantic projection is active, exact reconstruction and monotonicity remain valid, but unconditional Shapley equivalence is not claimed.
- Joint Optimization: The four design elements are jointly optimized with a class-weighted binary cross-entropy objective plus sparsity, semantic-budget, and residual-separation penalties.The regularization terms operationalize structural requirements rather than serving only as generic optimization devices.
- Prediction and Evaluation: The same hazard-to-probability mapping is used during training and inference, with the operating threshold locked before test evaluation.The threshold is selected from five-fold out-of-fold development predictions.
- Prediction and Explanation: A single forward pass yields prediction together with generic-safety, mechanism-level, and concept-level decomposition without separate post-hoc explanation.The decomposition reports shares of total hazard attributable to generic safety, mechanisms, and baseline risk.
- Artifact-Level Outcome: ChildRiskGuard is the retained outcome of AI4CDS resource construction and design search, with four elements addressing the paper’s three technical challenges.The documented search includes architectures, mechanisms, constraints, and response functions that were accepted, modified, reclassified, or rejected.
4. Evaluations
ChildRiskGuard is evaluated through preregistered-style, development-only selection, controlled ablations, benchmark comparisons, and explanation-faithfulness checks. It achieves F1 = 0.769 while providing mechanism- and concept-level attribution tied to its predictive computation.
- Evaluation design: Evaluation fixes rules before testing, isolates one focal quantity per ablation, compares benchmarks under common operating rules, and verifies AI-generated interpretations.These checks assess predictive utility, design-element and mechanism contributions, explanation faithfulness, robustness, and generalizability.
- Evaluation design: ChildRiskGuard is evaluated on 7,070 short-form videos using development-only fitting and threshold selection before final evaluation on an untouched test set.The data include 1,000 expert-annotated child-inappropriate videos and 6,070 additional manually reviewed videos; the test set contains 1,060 videos, including 150 positive cases.
- Predictive utility: F1 = 0.769 exceeds the strongest concept-bottleneck baseline, LM4CV at 0.758, while PCBM exhibits a different precision-recall trade-off rather than uniformly better performance.The benchmark families also include end-to-end video classifiers and large vision-language or language models.
- Interpretability: The artifact separates generic from child-specific risk, grounds mechanisms in developmental constructs, and integrates structured attribution into prediction without sample-level mechanism annotations.Compared prompting and fine-tuning baselines do not provide the same decomposition.
- Predictive utility: F1 = 0.769 exceeds frame-wise ResNet-18 at 0.711, while InternVL3 reaches 0.758 and fine-tuned LlavaGuard reaches 0.679.The results support incorporating domain-specific requirements into the predictive architecture rather than relying only on task-specific representation learning.
- Ablation analysis: The calibrated generic channel and monotone additive experts produce the largest ablation gains, increasing F1 by 0.085 and 0.077, respectively.The semantic budget adds 0.012, mechanism gating adds 0.006, and mechanisms M1, M2, and M3 add 0.035, 0.011, and 0.150.
- Interpretability: Across three cases, removing the focal mechanism reverses the final decision, supporting faithfulness between reported explanations and the model’s actual decision process.The examples show decision-critical contributions from dangerous action, adultized performance, and self-harm or depressive risk mechanisms.
5. Discussion and Conclusion
The discussion presents AI4CDS as a prescriptive framework for governing AI participation across computational design science while retaining human authority over scientific commitments. ChildRiskGuard supplies consequential feasibility evidence, but the framework and artifact remain bounded by their evaluated setting and require further examination across contexts.
- Knowledge abstraction: Phase 5 assembles requirements, decisions, evidence, corrections, and interventions, then tests candidate abstractions against evidence, case dependence, and contradictions before distilling transferable knowledge.The process yields methodological implications, governance principles, reproducibility guidance, domain implications, and boundary conditions.
- AI4CDS contribution: AI4CDS organizes AI participation across problem formulation, resource construction, design search, evaluation, and knowledge abstraction while researchers retain domain grounding, admissibility, verification, and scientific judgment.The framework makes roles, transitions, and responsibilities explicit for examination and refinement.
- Constitutive AI participation: Within fixed time and researcher constraints, AI broadened knowledge search, design alternatives, implementation, diagnosis, and cross-phase synthesis.The documented process indicates that AI materially expanded the breadth and depth of inquiry within the same research window.
- Constitutive AI participation: AI-generated outputs required verification, revision, or rejection before becoming scientific commitments, leaving researchers responsible for scientific commitments.The case therefore illustrates AI as expanding inquiry rather than determining scientific commitments.
- Governance: AI4CDS identifies constitutive-AI governance risks involving uncertain reliance, path-dependent commitment, and weakened provenance and accountability.Governance applies to transitions where provisional AI-supported outputs become or change scientific commitments.
- Governance: Auditability requires differentiated reproducibility: fixed procedures should be computationally reproducible, stochastic computations statistically reproducible, and AI-assisted search and diagnosis documented.These are presented as different forms of rigor rather than weaker levels.
- Artifact contribution: ChildRiskGuard translates audience-dependent safety and explanation faithfulness into an artifact that separates generic safety, child-specific mechanisms, and concept-level predictive explanations.Its evaluation shows substantial improvement over the reused generic-safety model and competitive performance relative to strong foundation-model alternatives.
- Domain implications: ChildRiskGuard’s outputs support human review, escalation, and auditing, but should not determine policy or intervention decisions.Mechanism-level decomposition distinguishes imitable dangerous actions, adultized performance, and self-harm or depressive evidence.
A.2 Phase 2: AI-Augmented Data & Resource Construction
Phase 2 converts AI-generated candidate resources into a validated, fixed resource base through researcher-led checks of construct validity, measurement behavior, structural role, and data-partition separation. The phase also encodes two forms of faithfulness in the final design.
- Candidate generation: AI proposes candidate representations, concepts, prompt wording, extractor families, structural assignments, and diagnostics, while researchers assess developmental fit, stability, specificity, and role.Assignments to mechanism blocks and activation families are fixed before risk-assessment parameters are fitted.
- Measurement stage: The fixed measurement stage produces a frozen general-safety reading from LlavaGuard and normalized concept measurements from frozen visual-language resources.The general-safety reading feeds a trainable calibration, while semantic concepts are scored separately.
- Mechanism resources: ChildRiskGuard assigns M1 to imitable dangerous action, M2 to adultized performance, and M3 to self-harm or depressive risk.M1 combines motion and pose with risky-action semantics; M2 combines skin exposure with suggestive context; M3 combines darkness, self-harm recall, and depressive affect.
- Mechanism resources: Activation is selective: M1 uses motion families, M2 uses a conjunctive skin-exposure and adultized-performance gate, and M3 is activated by darkness.Other assigned concepts affect mechanism intensity without independently opening its gate.
- Validation checks: Resource construction applies four checks: construct validity, measurement behavior, structural role, and separation of data partitions.These checks determine whether candidate resources are admissible and how they enter later computation.
- Resource refinement: Resource diagnostics can remove unsupported mechanisms, replace hard gates with differentiable family gates, and distinguish gate-opening evidence from intensity evidence.A dangerous-object mechanism was rejected because its measurements lacked specificity and became too sparse after filtering.
- Faithfulness: Faithfulness to model inputs is supported by fixed concept definitions, structural assignments, semantic caps, and mechanism-specific semantic budgets.These controls constrain how semantic evidence enters the child-specific pathway.
- Faithfulness: Faithfulness from evidence to prediction is supported by an anchored, monotone, additive hazard construction.The construction links evidence contributions to the predictive output and supports concept-level explanations.
A.3 Phase 3: AI-Generative Research Design Search & Refinement
Phase 3 turns domain challenges into testable design requirements and searches candidate mechanisms one component at a time. AI broadens alternatives, but researchers retain responsibility for admissibility, evaluation, critical review, and consequential design decisions.
- Design requirements: Phase 3 uses development data only for normalization, model selection, and threshold selection, pairing each retained design element with an isolating evaluation.These conditions define the admissible design space.
- Design requirements: The design space requires child-risk evidence to participate in prediction, developmental grounding for mechanism labels, no concept-level annotations, and faithful evidence-to-prediction mapping.These requirements constrain admissible artifact designs before candidate generation.
- AI-assisted search: AI broadens the candidate set with alternative model families, decompositions, response functions, constraints, and diagnostic tests.Researchers assess challenge fit, supervision and resource compatibility, and independent justification before accepting, modifying, reclassifying, or rejecting candidates.
- Design process: The recurring design process translates each challenge into a testable requirement, generates mechanisms one component at a time, and specifies evaluation before selection.This sequencing keeps iterations focused and makes candidate implications explicit.
- Critical review: AI-supported critiques help identify unsupported guarantees, incompatible requirements, and claims mismatched with the current architecture.Researchers then narrow claims, revise constructions, or reject components.
- Documentation: The design record documents each candidate’s origin, disposition, and decision basis, distinguishing theoretical judgments from measurement, resource, and development-data evidence.This record helps prevent rejected alternatives from being reconsidered without new evidence.
- Division of labor: AI contributed more to generating candidates than selecting the final design, because many candidates were revised or rejected after theoretical review, measurement analysis, or diagnostics.This division of labor is consistent with AI4CDS: AI broadens search while researchers remain responsible for admissibility.
A.4 Phase 4: AI-Enabled Multi-Faceted Evaluation
Phase 4 separated specification, execution, and interpretation so researchers could verify implementations, outputs, and claims before accepting results. AI supported implementation, diagnostics, benchmark comparisons, and candidate interpretations, while researchers retained evaluation and interpretive responsibility.
- Evaluation governance: The evaluation separated the designed method, implemented code, and reported claims to make discrepancies easier to identify.Researchers reviewed specifications before execution and checked values and artifacts before accepting interpretations.
- Evaluation governance: Researchers defined evaluation rules in advance, including data partitions, primary metrics, threshold selection, and comparison rules without using test labels for model selection.Operating thresholds were selected using development data held out from the corresponding model fit.
- Evaluation governance: Ablations changed one focal quantity at a time, reported checks on unauthorized quantities, and handled further corrections separately.This design was intended to attribute component effects to the intended change.
- Interpretation: AI-generated explanations were treated as proposals, and researchers could request diagnostics, narrow claims, or reject interpretations when evidence conflicted.Researchers examined conflicting splits, unsupported causal interpretations, and anomalies.
- Interpretation: Important corrections often arose at specification-to-execution or execution-to-interpretation handovers, especially for analyses that could make results appear stronger.These patterns support human review when computational outputs become experimental results or scientific claims.
- Knowledge abstraction: Phase 5 linked requirements, design decisions, corrections, and evaluation evidence before researchers judged which patterns warranted transferable design knowledge.AI organized records and proposed patterns, but researchers verified links, assessed scope, and rejected or narrowed unsupported abstractions.
B.2 Model Architecture and Training
ChildRiskGuard uses frozen measurements and structured, monotone risk components to separate generic safety from child-specific mechanisms. Training adds semantic-budget and residual-separation constraints, with the operating threshold selected from out-of-fold development predictions.
- Generic-safety channel: The generic-safety channel calibrates a frozen reading through eight learned knots with positive increments and a bounded terminal value.The first knot is zero, and the generic-hazard cap is H_G = 4.5.
- Mechanism structuring: Child-specific mechanisms combine fixed concept families through monotone responses, smooth family summaries, and conjunctive gates.M1 covers motion-related concepts, M2 gates skin exposure with adultized context, and M3 contains darkness.
- Semantic contribution budget: The model bounds semantic contributions using a per-concept cap, a soft budget penalty, and hard inference projection.The final model uses α_sem = 0.35 and H_sem = 2.5.
- Training: Only the risk-assessment stage is optimized; measurement components remain frozen during training.Trainable parameters include response coefficients, calibration coefficients, gate thresholds, and a nonnegative leak hazard.
- Calibration and decision rule: The operating threshold is selected by maximizing F1 on combined five-fold out-of-fold development predictions before final fitting and test evaluation.The selected threshold is η = 0.6123511.
- Training objective: Residual separation is a soft dependence regularizer rather than an independence guarantee.It discourages unnecessarily large or diffuse child-specific mechanism hazards.
B.3 Theoretical Results and Proofs
The theoretical results establish identifiable channel decomposition, monotonicity and faithfulness properties, and a conditional Shapley correspondence for ChildRiskGuard. The correspondence is limited when semantic projection creates coupling among concepts.
- Identifiability: The canonical cumulative-risk form is F(t) = 1 − exp(−t) after rescaling additive risk quantities.This follows from continuity and strict monotonicity under the additive decomposition conditions.
- Identifiability: Under the stated block-support and anchoring conditions, the additive generic and mechanism channel decomposition is identifiable.Every concept has one structural destination, and channel contributions cannot be arbitrarily reassigned.
- Faithfulness axioms: The child-specific pathway satisfies the four operational faithfulness axioms under the stated conditions.These include monotonicity, structural assignment, and exact hazard changes when mechanisms are removed.
- Conditional Shapley correspondence: When semantic projection is inactive, each concept’s marginal contribution is coalition-invariant conditional on the realized mechanism gate, yielding a Shapley interpretation.This correspondence does not treat gate-induced interactions as additive concept effects.
- Conditional Shapley correspondence: When semantic projection is active, the mechanism remains exactly reconstructable from reported contributions, but unconditional exact Shapley equivalence does not hold.The aggregate minimum couples semantic concepts.
C.1. Generalizability to an Additional Child-Specific Risk Mechanism
The generalizability analysis adds an independently grounded child-specific mechanism, M4, while retaining ChildRiskGuard’s existing architecture and M1–M3 specification. It tests architectural extensibility rather than completeness of the risk taxonomy.
- Generalizability analysis: Adding M4 tests whether ChildRiskGuard can accommodate a previously unmodeled child-specific risk mechanism.The analysis does not imply that four mechanisms exhaust relevant risks.
- M4 extension: M4 is defined from external substantive knowledge and implemented through a separate concept block and mechanism-specific evidence structure.It is introduced as an independently grounded mechanism alongside M1–M3.
- M4 extension: Extending ChildRiskGuard to M4 requires new M4-specific measurements but not redesign of the core risk-assessment architecture or fine-tuning of the frozen foundation model.The existing M1–M3 specification is retained while the risk-assessment stage is re-estimated.
C.2. Sensitivity to the Semantic Representation
The sensitivity analysis tests the semantic readout representation while holding the rest of ChildRiskGuard’s architecture unchanged. The reference representation performs best, making semantic measurement an important design component and boundary condition.
- F1 decreases from 0.769 with the reference representation to 0.731 with OpenCLIP and 0.709 with SigLIP.These correspond to drops of 0.038 and 0.060, respectively.
- The semantic representation aligned with the reused safety model provides a measurable performance advantage.
- Alternative encoders indicate that ChildRiskGuard is sensitive to the source of semantic measurement.
- Semantic representation is both an important design component and a boundary condition of the artifact.
D.1. Conceptual Basis and Theoretical Derivation
The theoretical derivation frames AI4CDS governance around consequential scientific transitions, where AI-supported outputs become researcher-authorized commitments. It distinguishes AI discretion from nondelegable human scientific authority and derives Graduated Trust, Reversibility, and Auditability.
- Scientific commitments: Scientific commitments are researcher-authorized decisions that constrain subsequent work or support a scientific claim.Examples include admitting resources, retaining design choices, fixing evaluation specifications, accepting evidence, and abstracting transferable claims.
- Consequential scientific transitions: AI4CDS treats consequential transitions as its primary governance unit because they affect scientific commitments.These effects can involve downstream propagation, evidentiary significance, claim significance, or correction costs.
- AI discretion and human authority: AI may generate candidate outputs, but human researchers retain authority over which outputs become scientific commitments.AI discretion covers intermediate actions within researcher-defined task boundaries, while scientific authority remains nondelegable.
- Graduated Trust: Graduated Trust calibrates AI discretion to task constraints, independent verification, and scientific consequence rather than research phase alone.Tightly specified computations may permit more discretion, whereas open-ended interpretation or abstraction may require stronger human authority.
- Reversibility: Reversibility requires reopening earlier commitments when credible later evidence materially challenges their assumptions, requirements, measurements, or justification.Researchers then decide whether to retain, revise, replace, or reject the commitment.
- Auditability: Auditability requires an inspectable record of consequential transitions sufficient to reconstruct AI contributions, researcher decisions, supporting evidence, and consequences.Documentation should scale with scientific consequence rather than recording every routine interaction equally.
D.3. Differentiated Reproducibility as an Auditability Standard
Differentiated reproducibility matches verification to the computational and epistemic character of each AI-enabled CDS activity. This avoids demanding exact reproduction from stochastic or generative work while preserving rigorous independent evaluation.
- Rationale: Using one reproduction standard across AI-enabled CDS would impose infeasible exact-reproduction demands on stochastic work or insufficient evidence on directly rerunnable computation.
- Verification standards: Differentiated reproducibility matches each research activity to an appropriate form of independent verification.The standard depends on determinism, stochasticity, process dependence, and the nature of the claim.
- Verification standards: Deterministic computation requires reproducible execution, stochastic computation requires stable results, AI-assisted search requires reconstructable decisions, and formal claims require valid derivations.
- Procedural reconstructability: Procedural reconstructability does not require identical generative outputs; it requires provenance sufficient to evaluate how alternatives became authorized commitments.
- Proportionality: Governance effort should increase with scientific consequence, verification difficulty, and propagation potential.Routine assistance may receive lightweight documentation, while decisions affecting evidence, artifacts, results, or transferable claims require stronger records.
- Governance priority: At consequential scientific transitions, validity and accountability take precedence over efficiency or convenience.The framework combines discretion, reopenable commitments, and traceable transitions within explicit boundary conditions.