Source-linked AI summary

From Subjective Judgments to Auditable Standards:Protocol-Guided AI Auditing of Website Redundancy

Ge Kong, Yongtong Cao

arXiv:2608.21476v1cs.SEcs.CV

TL;DR

Website redundancy changes meaning across users, tasks, and failure scenarios, so a single score is insufficient for consistent evaluation. CORA factorizes redundancy into load, normal-use tax, and failure-domain reserve, then audits model evidence and releases scores only when checks pass. On a controlled testbed it separated reserve from tax and predicted perturbed success better than scalar-load baselines, while withholding scores from both model instruments that failed release requirements.

  • Problem

    Website redundancy depends on users, tasks, failure scenarios, and repetition forms, while repeatability alone does not establish validity.

  • Method

    CORA measures load, normal-use tax, and failure-domain reserve separately, using versioned model annotations, retained audit artifacts, typed validation, controlled interventions, and conjunctive release gates.

  • Results

    On a transparent mechanistic testbed, CORA separated reserve from normal-use tax and predicted perturbed success more accurately than scalar-load baselines; both model instruments were withheld after failing release checks.

  • Takeaways & Limitations

    CORA provides an auditable candidate procedure for the controlled benchmark by recording evidence and explaining score release or refusal.

  • Takeaways & Limitations

    Human agreement, AI-versus-human accuracy, and validation on independent production sites remain untested, so CORA is not yet a general standard.

Abstract

from arXiv · show

Website redundancy does not have a single fixed meaning. The same repeated element may distract during one task and provide backup during another. We introduce CORA (Counterfactual, Observable Redundancy Audit), which measures repetition load, normal-use tax, and failure-domain recovery reserve separately. Each run retains screenshots, stable element identities, and task traces. A versioned vision-language model proposes the annotations. Typed validation and release checks then determine whether a calibrated dimension can be reported; failed or malformed outputs stay in the fixed denominator. On a transparent mechanistic testbed, the factorized CORA representation separated reserve from normal-use tax and predicted perturbed success more accurately than scalar-load baselines. The model studies then showed why repeatability is not enough: two small local vision-language models produced recurring outputs, but neither instrument met all release requirements. CORA therefore withheld automated scores from both instruments while retaining the raw responses and failure records. Separate checker fixtures confirmed that the typed validator and hardened release gates implement their specifications; these tests do not establish semantic grounding or accuracy on production sites. Taken together, the results position CORA as an auditable candidate procedure for the controlled benchmark studied here rather than a general standard. Human agreement, AI-versus-human accuracy, and validation on independent production sites remain open empirical questions.

1 Introduction

CORA treats website redundancy as a task- and failure-dependent measurement problem rather than a single subjective property. It separates repetition load, normal-use tax, and recovery reserve, then uses auditable checks to release or withhold scores.

  • Website redundancy varies with the user, task, failure scenario, and form of repetition, making consistent measurement difficult.
  • CORA separates repeated structure, ordinary-use cost, and post-failure usefulness instead of collapsing them into one clutter score.A confirmation step can raise tax without adding reserve, whereas an alternate route adds reserve only when it survives a different failure domain.
  • The audit contract fixes the task, page, viewport, browser, model revision, prompt, and success condition while retaining screenshots, stable element identities, and task traces.The model proposes semantic annotations and typed citations; deterministic checks measure routes and task states.
  • Two model studies produced recurring outputs but failed different release requirements, so CORA withheld both automated scores while retaining their records.Qwen gained a small amount of repetition ordering after an output-contract change but failed grounding and coverage; SmolVLM2 echoed traces instead of returning the requested schema.
  • A conjunctive release gate requires parsing, evidence count, repetition ordering, grounding, target, qualified replay, and non-degeneracy checks to pass before reporting a controlled dimension.Uncontrolled reserve and clarity fields remain ineligible for release.
  • The workflow was evaluated through a transparent mechanistic testbed, four seed replications, a 77-call discovery study, a 336-call calibration repair, and a 210-call second-model replication.The studies separated repeatability, discrimination, grounding, and safe refusal across model configurations.

2 Related work

Related work distinguishes usability dimensions, evaluator effects, visual complexity, grounding, attribution, and behavioral testing. CORA builds on these distinctions by treating redundancy and evidence support as separate, testable measurement conditions.

  • Visual complexity can be measured and related to search or aesthetic judgments, but it is not synonymous with redundancy.Redundancy effects depend on the user and representation rather than being uniformly beneficial.
  • Prior usability research reports variation in operationalization and generally weak associations among effectiveness, efficiency, satisfaction, and subjective versus objective measures.Agreement measures require interpretation in light of rating design.
  • Automated usability evaluation spans methods including visual clutter, visual learnability, learned interaction flows, structured UX instruments, and aggregated agent traces.
  • Language and vision-language models show evaluator effects and dimension-dependent agreement, while UI benchmarks report limited understanding of behavior-changing interfaces.
  • GUI grounding is distinct from high-level judgment, motivating separate checks for structure, visible support, and final judgment.Prior findings do not validate CORA’s specific region-overlap threshold or evidence-count rule.
  • Attributed-generation research separates answer correctness from citation support, paralleling CORA’s distinction between evidence correctness and output-level coverage.Text-attribution studies do not validate CORA’s specific interface-measurement rules.
  • Behavioral test suites, versioned artifacts, explicit reporting, and selective classification motivate CORA’s failure-mode testing, traceability, and abstention design.

3 Measurement protocol

CORA operationalizes redundancy as separate load, normal-use tax, and failure-domain recovery reserve, then releases scores only for covered dimensions that pass conjunctive checks. The protocol preserves observable evidence and fixed-denominator failures while withholding unqualified automated scores.

  • Measurement dimensions: CORA separates repeated structure, ordinary-use work, and useful alternatives that survive a matching failure.Load covers visual, information, and interaction repetition; tax covers normal-use cost; reserve covers surviving cues, landmarks, or routes.
  • Task-conditioned measurement: A fixed observation contract defines the interface, task, evidence, and machine-verifiable success states for each audit.The interface is modeled as a state-action system with states, actions, transitions, and observable evidence.
  • Evidence grounding: Evidence citations must use valid catalog IDs, match visible text exactly, or meet the typed region-overlap rule.Each call requires 2–4 evidence items, and pooled grounding cannot replace output-level coverage.
  • Fixed-denominator evaluation: Every scheduled output remains in the denominator, so parse failures are counted as incorrect rather than retried, repaired, or imputed.Conditional Kendall τ is reported only when both vectors vary, while Study A intervals describe six overlapping held-out folds.
  • Dimension-scoped release: The conjunctive release rule reports only covered dimensions when parse, evidence, ordering, grounding, target, replay, and non-degeneracy checks pass.Otherwise CORA returns WITHHOLD_AUTOMATED_SCORES while retaining raw evidence and diagnostic-only reserve and clarity.

4 Experiments

The experiments use parameterized rendered pages and frozen model configurations to test identifiability, auditability, reproducibility, and refusal under controlled benchmark variation. Studies span mechanistic load-matched testing, local-VLM evaluations, typed-grounding checks, and end-to-end release-gate fixtures.

  • Study A: exactly load-matched identifiability: Study A controls six task configurations, three redundancy types, load budgets, reserve-tax allocations, perturbations, noise levels, and 1,555,200 episodes.All nine type-by-budget strata have equal scalar load and include all four reserve-tax compositions.
  • Study A: exactly load-matched identifiability: Study A regresses retention on matching reserve and predicts perturbed success with leave-one-family-out evaluation against scalar baselines.The strongest scalar baseline is a 500-tree random forest without hyperparameter search.
  • Study A: exactly load-matched identifiability: The follow-up reuses the primary seed plus four unchanged master seeds, reporting descriptive ranges and standard deviations rather than population inference.The five full runs contain 7,776,000 episodes.
  • Model studies: Study B evaluates Qwen2.5-VL-3B-Instruct on seven e-commerce-page variants using screenshot-only, evidence-only, and fused contracts across 77 calls.Study C evaluates two Qwen arms on 42 pages across six task/content domains sharing one HTML/CSS layout, with 336 total calls.
  • Model studies: Study D runs unchanged C-ORD with SmolVLM2-2.2B-Instruct on the same 42 pages, using three stochastic seeds and two deterministic replays for 210 calls.Only model-specific loading and chat templating change; checkpoint, prompt, stimuli, runner, thresholds, and metric concepts remain locked.
  • Checker fixtures: Separate suites test typed citation validation, complete release-gate processing, and output-level evidence cardinality on valid and invalid panels.The end-to-end suite processes 672 strings, while the evidence-cardinality suite processes 336 raw strings.

5 Results

The controlled benchmark found that CORA’s factorized representation separated recovery reserve from normal-use tax and predicted perturbed success better than a scalar-load baseline. Model studies showed that repeatability did not suffice for qualification: both instruments failed release requirements, while checker fixtures confirmed specification conformance without establishing real-world validity.

  • CORA separates reserve from tax within the planted mechanism: Reserve had a positive retention slope in 319/324 analyzable strata (98.5%), with median within-stratum Spearman ρ of 0.894.All six family median slopes were positive, with exact two-sided sign test p = 0.03125 over six fixed configurations.
  • CORA separates reserve from tax within the planted mechanism: High-tax profiles added 1.165 nominal actions on average despite strictly matched reserve, with positive differences in all 324 descriptive pairs.All six family means were positive, supporting separation of normal-use tax from recovery reserve in the mechanistic testbed.
  • CORA separates reserve from tax within the planted mechanism: Full CORA achieved LOFO RMSE 0.0222 and R2 = 0.9559, versus RMSE 0.0469 and R2 = 0.8030 for the random-forest scalar.The relative RMSE reduction was 52.8%, and all six overlapping paired folds favored CORA.
  • Qualification distinguishes repeatability from discrimination: Study B produced highly repeatable outputs that copied the example’s 0/100 pattern, but controlled ordering was zero.All 7/7 deterministic replay pairs were byte-identical, while 77/77 raw outputs copied the pattern.
  • Qualification evaluates the complete output contract: After a complete Qwen output-contract change, strict ordering rose to 11/108 (10.2%), but typed grounding remained 0/358 and no score was eligible for release.The combined changes show sensitivity to the complete contract rather than isolating numeric anchors.
  • The same qualification gate supports cross-model comparison: SmolVLM2 generated extractable JSON in 82/126 calls but no schema-valid ordinal object, so the unchanged gate withheld the second instrument.Its outputs copied deterministic observation traces, and protocol-qualified replay was 0/42.
  • Executable conformance tests verify the release path: Typed v2 achieved F1 1.00 on 504 citation cases, while the legacy validator achieved F1 0.50 with 126 false accepts and 126 false rejects.The end-to-end full gate released all 16 valid panels and withheld all 16 invalid panels; these fixtures test specification conformance, not production error rates.
  • Executable conformance tests verify the release path: Requiring 2–4 evidence items per scheduled output withheld all 8 invalid panels while releasing all 8 valid panels in the cardinality hardening suite.The suite covered enumerated cardinality faults rather than a broader fault family.

6 Discussion

CORA frames procedural objectivity as inspectable measurement practice rather than evidence that a model is unbiased or correct. Its records preserve useful failure information, but broader standardization requires independent calibration and validation beyond this controlled benchmark.

  • What repeated runs can establish: Procedural objectivity means another researcher can inspect what the model saw, which checks failed, and why a score was released or withheld.The term applies to the audit procedure and does not establish model unbiasedness or correctness.
  • What repeated runs can establish: Withheld runs retain screenshots, page catalogs, task traces, raw responses, and failed checks, exposing why stable outputs did not qualify as website-quality measurements.Thus abstention preserves an auditable record rather than discarding the run.
  • Interpreting the contract comparison: The Qwen contract comparison cannot attribute the ordering gain uniquely to numeric anchors because scale, schema, evidence format, and examples changed together.Ordering improved while schema validity, target accuracy, grounding, and non-degeneracy still failed.
  • Interpreting the contract comparison: Under the same ordinal prompt, Qwen returned partly parseable but poorly grounded judgments, whereas SmolVLM2 copied observation traces without schema compliance.Separate failure records made the model-specific violation patterns comparable.
  • Scope beyond the controlled benchmark: Before serving as a standard, CORA needs external calibration across sites, experts, users, languages, devices, and accessibility conditions.The paper also requires fixed tasks and rendering, inspectable evidence for every output, and dimension-specific controlled interventions.
  • Scope beyond the controlled benchmark: The load-tax-reserve hypotheses remain grounded in the mechanistic benchmark and still require tests on real interfaces.Examples include testing cue independence for information repetition and requiring alternate routes to survive a different failure domain.

7 Limitations

The studies constrain CORA’s evidence to planted mechanisms, controlled benchmark variation, and small local-model settings rather than production websites or human judgments. Several protocol and coverage limitations further restrict generalization across layouts, failure domains, and access needs.

  • No human participants or production sites were used, so inter-rater reliability, expert agreement, perceived clarity, task time, preference, and observed recovery remain unmeasured.
  • Study A used six parameterized families from one planted mechanism; five seeds measure seed sensitivity, while intervals describe dispersion rather than population uncertainty.The minimum two-sided sign-test value under agreement across all six families is p = 0.03125.
  • Studies C and D test contract response under controlled benchmark variation, not open-world visual understanding, using one shared layout and deterministic structural counts.
  • The oracle controls visual, information, and interaction repetition but leaves AI-estimated reserve and clarity fields diagnostic rather than validated dimensions.Validator F1 1.00 and 0/16 false releases apply only to frozen specification fixtures and do not establish semantic relevance or production error rates.
  • The Study C protocol originally claimed randomized layouts, but the rendered artifact used one shared layout template; conservative corrections were recorded without changing raw outputs.
  • The evaluation excludes multilingual, mobile, accessibility-specific, and unconstrained computer-use-agent settings, each of which requires explicit failure domains and outcomes.

8 Ethics, data governance, and AI assistance

The studies used rendered benchmark pages, generated trajectories, and local checkpoints without personal data. Human or production studies would require consent, privacy protections, research-data treatment for interaction traces, and separation of observations from model inference.

  • The experiments used programmatically rendered benchmark pages, generated task trajectories, and local open model checkpoints, collecting no personal data.
  • Future human or production studies would require consent, masking typed payloads, no retention of personal form content, and treating cursor and interaction traces as research data.
  • Reports should distinguish deterministic observations from model inference and state calibration status.
  • Generative AI assisted development, orchestration, analysis checks, and drafting, while numerical claims were checked against machine-readable artifacts or deterministic validators.Tool assistance was not treated as authorship or independent evidence.

9 Reproducibility and data availability

The local artifact preserves protocols, outputs, corrections, manifests, and conformance tests, supporting byte-level regeneration and audit tracing. External reproduction remains limited because no persistent public repository exists and full checkpoints and episode tables are unavailable in the source package.

  • The artifact contains frozen protocols, lock files, raw outputs, corrected analyzers, manifests, hashes, metrics, figure scripts, citation records, and conformance tests.
  • Study A’s eight primary artifacts were regenerated byte-for-byte at the same seed, with four additional full master seeds stored alongside manifests and aggregate metrics.
  • Study C preserves all 336 raw calls, while post hoc conservative replay and evidence-coverage corrections bind old and new analyses.
  • No persistent public repository URL had been assigned at the preprint date, so public code and data release is still needed for external reproduction.
  • The arXiv source package includes the manuscript and vector figures but not multi-gigabyte model checkpoints or the full episode table.

10 Conclusion

CORA frames website-redundancy judgment as a testable measurement problem and separates its dimensions while auditing model outputs. In the benchmark, it recovered planted behavior, withheld both VLM scores after distinct failures, and remains an executable proposal requiring external testing rather than a general standard.

  • CORA separates load, tax, and reserve, recovering planted behavior across five master seeds in Study A.
  • Both VLMs produced repeatable outputs but failed release checks: Qwen failed grounding and coverage, while SmolVLM2 failed to return the requested schema.
  • The checker recorded both failure patterns and withheld both automated scores, preserving the audit distinction between repeatability and validity.
  • CORA’s main contribution is a reproducible account of why a score was released or withheld, including page evidence and failed checks.
  • CORA is currently an executable measurement proposal for external testing, requiring independent sites and human judgments across languages, devices, and access needs before becoming a general standard.

A Study A mechanistic specification

Study A specifies controlled profiles that vary reserve and tax across visual, information, and interaction families, while holding other dimensions medium. Its transparent mechanistic testbed links signal, routes, ambiguity, noise, and distraction to correct-action probability, but is not a model of human behavior.

  • Profile construction: Profiles allocate a fixed budget across reserve, tax, and neutral units, with family-specific constructions for visual, information, and interaction redundancy.Low/high reserve receive B/4 and B/2 units; low/high tax receive 0 and B/4 units, respectively.
  • Mechanistic specification: The mechanistic specification models correct-action probability from available signal, routes, family ambiguity, observation noise, and distraction burden.The displayed expression assigns positive coefficients to signal and routes and a negative coefficient to distraction burden.
  • Mechanistic specification: Tax units produce stochastic nominal actions with probability min(0.85, 0.50 + 0.35n), while matching structural-loss probability is 0.45.The mechanism is transparent by design and should not be interpreted as a model of human behavior.
  • Evaluation: Study A evaluates perturbed-success prediction with leave-one-family-out testing and scalar baselines that lack profile metadata and component counts.The supplied table caption identifies this evaluation as Study A leave-one-family-out perturbed-success prediction.
Loading 2608.21476v1…