Source-linked AI summary
Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress
Nataliya Shakhovska, Ivan Izonin, Stergios-Aristoteles Mitoulis
TL;DR
Prediction-side certificates may not establish trustworthy AI because reliable and compromised models can look identical in prediction behaviour. The paper proves this separation and introduces a competence envelope combining prediction and explanation certification, showing that decision-mechanism evidence is also required.
Problem
Existing trustworthy-AI methods do not certify the joint region where both a model’s prediction and explanation can be trusted.
Method
The paper proves a prediction–explanation separation theorem and defines a competence envelope combining predictive reliability with bounded explanation fidelity.
Results
A reliable and compromised model can match every prediction-side certificate while differing arbitrarily in explanation fidelity and deployment behaviour.
Takeaways & Limitations
Trustworthiness certification for this failure class requires evidence about the model’s decision mechanism in addition to its predictions.
Takeaways & Limitations
The empirical study exercises three of four competence-envelope axes and makes no claim of universality.
Abstract
from arXiv · showhide
Artificial intelligence systems increasingly make consequential judgments - which patient is deteriorating, which building is safe to enter, whether an image is authentic and are trusted on the strength of how accurately and confidently they predict. The safeguards that certify them are correspondingly prediction-based: accuracy, calibration and conformal coverage all measure how well a model performs. Whether such checks are sufficient to establish model trustworthiness has remained unclear. Here we prove that they cannot. We establish a separation theorem showing that a reliable model and a compromised one can be identical under every prediction-side certificate, including accuracy, calibration and coverage, yet differ arbitrarily in explanation fidelity and deployment behaviour. Detecting this failure requires access to the model's decision mechanism in addition to its predictions. We introduce the competence envelope as an operational framework that combines prediction and explanation certification into a single deployable criterion. Across diverse datasets and model classes, the proposed framework reveals failure modes that prediction-side certification alone does not capture. Certification against failures that are invisible in prediction behaviour therefore requires evidence about the model's decision mechanism as well as its outputs.
Introduction
Prediction-side certificates can leave compromised systems indistinguishable from reliable ones while explanations remain untrustworthy and deployment behavior fails under shift. The paper therefore argues that trustworthy certification must jointly certify prediction reliability and explanation fidelity across operating conditions.
- Motivation: Automated systems increasingly shape consequential judgments, yet their controlled-evaluation accuracy depends on assumptions that may fail in deployment.Examples include deterioration assessment, building re-entry safety, and authenticity judgments.
- Problem: The central danger is silent failure: systems can leave their competence conditions while continuing to issue confident answers with plausible-looking reasons.Machine-reliant cues may differ from cues people would trust.
- Core result: The separation theorem shows that reliable and compromised models can match every prediction-side certificate while differing arbitrarily in explanation fidelity and behavior under shift.The listed certificates include conformal coverage, accuracy, calibration, and confidence.
- Core result: A matching lower bound shows that prediction-only monitors cannot detect this failure above chance, so certification requires reading the model’s structure and reasons.The paper states that certifying explanations is necessary for certifying against this failure class.
- Proposed framework: The competence envelope couples certified explanation fidelity with prediction reliability over operating conditions, addressing the joint region that existing guarantees do not characterize.The paper identifies this as a structural gap spanning shift, scarcity, robustness, abstention, and explanation assurance.
Results … An empirical competence envelope
The paper defines competence as a jointly certified region over measurable operating conditions, requiring both predictive reliability and explanation fidelity. Theorems and experiments show that prediction-side certificates can miss structural compromise, while explanation certification and compound-stress analysis reveal silent failures and envelope contraction.
- A measurable space of operating conditions: The competence envelope C(f,E) is the region of operating conditions where predictive reliability and explanation fidelity are jointly certified.Operating conditions span drift, scarcity, contamination, and resource degradation, represented as monotone stress signals in Ω.
- Why prediction certification cannot replace explanation certification: Theorem 1 proves that two models can match every prediction-side certificate while differing arbitrarily in explanation fidelity and deployment accuracy.For any fidelity gap β ∈ (0,1) and accuracy margin γ ∈ (0, ½), the compromised model can satisfy F(f′) ≤ 1 − β and have deployment accuracy at least γ lower.
- Why prediction certification cannot replace explanation certification: Prediction-law monitors have no detection advantage against the constructed compromise, so detecting it requires a structural functional such as a sensitivity map.The result covers accuracy, calibration, coverage, Brier score, AUC, confidence, post-hoc combinations, and split-conformal coverage.
- Certifying explanations, not only predictions: Under monotone stress signals, the jointly certified envelope is non-empty and star-shaped, with its radial boundary set by whichever certificate fails first.Proposition 1 expresses this geometry as K(α,β) = { ω : C(ω) ≥ 1−α and F(ω) ≥ 1−β } and ∂K(u) = min(tC, tF).
- Certifying explanations, not only predictions: Explanation certification combines faithfulness and stability functionals, conformalised over Ω to produce a distribution-free explanation-competence region alongside predictive reliability certification.Faithfulness uses perturbation-based infidelity and deletion/insertion agreement, while stability uses local explanation-map behaviour and class-specific bounds where available.
- Compound stress and the contraction of competence: Compound stress is modelled through interacting axes, with contraction measured by interaction terms such as g(σ,ρ) = g₀ − aσ − bρ − c·σρ.Positive c indicates that the envelope shrinks faster than either stress axis alone predicts, while c ≈ 0 indicates independent effects.
- An empirical competence envelope: On Telegram and Reddit data, the framework replicated certificate-specific degradation and identified silent failures that prediction-side metrics did not yet reveal.On Telegram, scarcity withheld the stability certificate at n = 300 with profile distance 0.33 while coverage was 0.91 and accuracy 0.72; on Reddit, coverage fell from 0.94 in 2019 to 0.85 by 2023 at n = 2,400, while stability rose from 0.12 to 0.32.
The envelope pattern persists across the tested architectures … Competence-gated fusion makes multimodal systems fail-safe
The competence envelope persists across architectures, foundation-model representations, datasets, and a language model, while label-free monitoring, resource-stress testing, and competence-gated fusion expose and mitigate failures that prediction certificates alone can miss.
- The envelope pattern persists across the tested architectures: Across five architectures, conformal coverage degrades under temporal drift while explanation-stability drift rises with scarce training data, demonstrating an architecture-independent envelope.The tested classes are logistic regression, MLP, GBDT, XGBoost, and LightGBM.
- Foundation-model backbones: The same two-certificate signature holds on frozen DistilBERT-multilingual, XLM-R, and multilingual MiniLM representations across Telegram and Reddit.Each backbone uses a certified logistic probe under the identical envelope protocol.
- Cross-dataset generality: Across eight text-classification benchmarks, explanation-stability degrades under training scarcity and coverage degrades with increasing covariate drift.The benchmarks span news, reviews, and social media tasks using DistilBERT-multilingual embeddings.
- Certifying a large language model’s explanations: On LoRA-adapted Qwen2.5-1.5B, coverage falls from 0.886 to 0.811 over five months while accuracy falls from 0.70 to 0.59.The acceptance boundary at 0.88 admits the first months and excludes later ones, matching the linear-model envelope boundary.
- Certifying a large language model’s explanations: Clean reference inputs that never activate a dormant trigger leave both prediction certification and explanation auditing insensitive to the planted backdoor.The language-model explanation-stability trend is also descriptive because integrated-gradient profiles are more variable under limited data.
- A label-free competence monitor anticipates silent failure: A label-free competence monitor combines input drift, attribution stability, and conformal uncertainty to anticipate silent failure without test labels.These signals are computed from an unlabelled discriminator, incoming attribution profiles, and conformal non-conformity.
- Resource degradation: the fourth independent axis: Pruning p = 0.70 drives S to 0.26 while coverage remains 0.88, and pruning p = 0.85 with cal_n = 20 drives coverage to 0.84.Reducing acoustic-visual bandwidth from full bandwidth to 15% raises deployment error from 0.43 to 0.51; poisoning raises S from 0.10 to 0.18 with r = 0.79.
- Competence-gated fusion makes multimodal systems fail-safe: Competence-gated fusion lowers failure error to 0.31 versus 0.45 on Reddit and 0.24 versus 0.33 when the tabular channel fails, while prescribing abstention when no competent channel remains.The routing principle is to use still-competent channels and refuse when none can safely support fusion.
Discussion
The discussion frames competence as an operational, real-time property requiring joint prediction and explanation certification under compound stress. It extends this framework toward resilience, consequential deployment, regulatory tooling, and stress-based benchmarking while acknowledging scope limits.
- Operational competence: An online competence signal should anticipate envelope departure, adapt explanations, and gate predictions among predicting, abstaining, deferring, or fallback actions.The signal is intended as a leading indicator rather than a post-hoc alarm, with thresholds set by the cost of silent failure.
- Operational competence: valid ⇔ covα(ω) ≥ 1−α ∧ Fα(ω) ≥ τ_F ∧ Sα(ω) ≤ L.The certificate also records envelope membership, predictive coverage, competence-signal calibration, fidelity, stability, and the prescribed action; failed validity requires a non-predicting response.
- Applications and stakes: Post-crisis reconstruction is presented as a severe competence test because compound stress, scarce data, consequential decisions, and auditability converge at scale.The discussion cites Ukraine’s projected reconstruction and recovery cost at almost US$588 billion, with direct physical damage exceeding US$195 billion.
- Benchmarking and regulation: The proposed programme calls for stress-organised benchmarks and resilience certificates that document coverage, calibration, fidelity, stability, and envelope membership for regulatory oversight.The framework is positioned as regulator-facing tooling supporting risk management, robustness, post-market monitoring, human oversight, documentation, and demonstrated functionality.
- Conclusion and limits: The central conclusion is that prediction-side evidence cannot establish a trustworthy decision mechanism, while the explanation certificate has its clearest interpretation for linear and tree-ensemble models.Evaluation of token-level integrated-gradient explanations in a LoRA-adapted autoregressive language model supports observability beyond frozen-representation probes but does not establish broader claims.
Methods
The methods formalize why prediction-side information cannot identify a model’s decision mechanism, then operationalize joint prediction and explanation certification across compound operating stresses, architectures, modalities, and domains.
- Operating-condition space: The competence envelope instruments drift, scarcity, adversarial contamination, and resource degradation using estimable stress signals and grids their combinations in the experiments.Signals include population-stability, Kolmogorov–Smirnov, maximum-mean-discrepancy, effective-sample-size, density, label-noise, shift, latency, memory, and energy measures.
- Information-theoretic formulation: Prediction-side information is defined as the joint law of inputs, outcomes, and predictions, so calibration, conformal prediction, uncertainty quantification, and selective prediction form one informational class.The framework contrasts this with structural information about the decision mechanism.
- Separation theorem: The non-identifiability lemma constructs models with identical prediction-side information but different structural information, while Theorem 1 makes the separation quantitative for fidelity gap β and accuracy margin γ.The results are proved constructively, with full proofs in the Supplementary Information.
- Empirical evaluation: Experiments span five model classes, a genuine Qwen2.5-1.5B language model adapted with LoRA, a clinical gradient-boosted-tree contrast, multimodal systems, and a 17-bridge infrastructure re-analysis.The infrastructure case additionally requires adequate LoK grades, deferring Low-LoK assets.
- Prediction certification: Predictive reliability uses split-conformal prediction with nominal target coverage 1−α = 0.90 and an operational envelope threshold of coverage ≥ 0.88.The quantile is set on an anchor-period calibration split, and the nominal target and acceptance threshold are kept distinct.
- Explanation certification: Explanation certification measures stability drift S = 1 − cos(φ, φ0), reserving fidelity F for the theorem’s conceptual structural-attribution quantity.Attribution profiles are instantiated per model class and compared with a large-sample or clean reference profile.
Code and data availability
The study uses publicly available datasets and open-source software, with complete reproducibility code released under the MIT License and archived at Zenodo.
- Data availability: All analysed datasets are publicly available, including Telegram messages for Domain A and climate-related Reddit posts for Domain B.Domain A includes 30,066 messages from eleven Telegram channels, while Domain B includes 36,642 Reddit posts available from the Harvard Dataverse.
- Code availability: The analyses use publicly available open-source software, including scikit-learn, SHAP, PyTorch and Hugging Face Transformers.These tools support the reported experiments and analyses.
- Code availability: Complete source code for reproducing every experiment, figure and numerical result is publicly available under the MIT License.The repository is available at https://github.com/natalya233/machine-competence-envelope.
- Code availability: The reproducibility code is also archived at Zenodo under doi:10.5281/zenodo.20771265.The archive provides a persistent record of the released source code.
Ethics statement
The study relied exclusively on publicly available datasets, involved no private communications, personally identifiable information, or human-subject interventions, and followed the original data sources’ terms of use.
- The study used only publicly available datasets.
- No private communications, personally identifiable information, or human-subject interventions were involved.
- All analyses followed the terms of use of the original data sources.
Supplementary Methods
The supplementary methods define shared representations, operating-condition axes, and a competence envelope combining predictive coverage with explanation stability. They also specify compound-stress error fitting and evaluation across Telegram, Reddit climate, and clinical tabular settings.
- Representation and model: Both domains use a shared 6,000-feature character word-boundary TF-IDF representation and an L2-regularised logistic classifier with C = 4.The representation is fitted once on a reference sample, enabling comparable feature spaces and attribution profiles across operating points.
- Operating-condition axes: Operating conditions vary temporal drift, training-set scarcity n, and adversarial label contamination ρ, while resource degradation is not exercised.Drift is measured by δ = |2(AUC−½)| from a held-out period discriminator; models train on January for Domain A and 2019 for Domain B.
- Certificates: The competence envelope requires split-conformal coverage ≥ 0.88 and explanation-stability drift S ≤ 0.20.Coverage targets 0.90 with α = 0.10, while stability is cosine drift in the global attribution profile relative to an in-distribution reference.
- Compound-stress analysis: Under compound scarcity × contamination stress, prediction error follows g(σ,ρ) = e₀ + aσ + bρ + c·σρ, with c estimated by least squares and bootstrap 95% confidence intervals.The clinical contrast repeats the grid on the Wisconsin Diagnostic Breast Cancer dataset using a gradient-boosted tree ensemble; Table S5 reports c ≈ 0 in all three domains.
- Evaluation grids: Supplementary grids evaluate coverage, explanation-stability drift S, accuracy, and confident-but-wrong rate across drift and scarcity in Telegram and Reddit climate domains.The Telegram grid averages over six seeds, and the Reddit climate grid replicates the two-certificate structure.
Supplementary results: label-free competence monitor
The label-free competence monitor combines input drift, explanation stability, and predictive uncertainty without test labels, then evaluates how these signals track held-out true error across deployment periods. Its certificate-gated abstention and joint certification analyses assess risk–coverage and show complementary predictive and explanatory margins.
- Construction: The monitor combines input drift, explanation stability, and predictive uncertainty, all computable without test labels.The signals are held-out discriminator drift, attribution-profile drift, and mean conformal non-conformity or 1 − max softmax probability.
- Construction: A model is trained once on an anchor period and carried forward across later periods, with true error held out for evaluation only.The anchor sets contain n = 4,000 examples for Domain A in January and Domain B in 2019.
- Monitoring results: The label-free monitor tracks the anchored model’s held-out true error across deployment periods, with Spearman ρ = 0.60 on Telegram and ρ = 0.90 on Reddit.Supplementary Table S6 reports these correlations using mean non-conformity without test labels.
- Abstention results: Certificate-gated abstention evaluates error on answered inputs as coverage is reduced by abstaining on the least-certifiable inputs.The analysis uses the area under the risk–coverage curve (AURC).
- Joint certification: The joint certificate matches or beats either certificate side alone, while the two margins are near-uncorrelated and therefore complementary.Supplementary Table S8 predicts true error from certificate signals using R² and reports cross-certificate margin correlation.
Supplementary results: adversarial capture and multimodal fusion
Supplementary results show that adversarial capture can remain invisible to validation accuracy and conformal coverage, while explanation-stability drift detects it. Competence-gated multimodal fusion and level-of-knowledge certificates also mitigate failures that confidence or naïve thresholds miss.
- Adversarial capture: As attack strength ρ rises, attack success increases while validation accuracy and conformal coverage remain green; explanation-stability drift S tracks the capture.This silent data-poisoning attack targets the Telegram amplification task.
- Multimodal fusion: Under modality-specific failure, competence-gated fusion tracks the single-best-modality oracle, whereas confidence-gated fusion matches naïve fusion because degraded channels remain confident.The comparison covers three systems under clean and modality-specific failure conditions.
- Certificate-gated damage assessment: For 17 Irpin bridges, a naïve coherent-change threshold flags 10 assets, while the level-of-knowledge certificate defers 2 confident misreads and identifies 1 high-damage verdict supported only by medium-reliability data.The certificate therefore distinguishes unreliable data from high-damage findings with limited reliability.
Supplementary Theory: proofs
The proofs establish that prediction-side certificates can be identical for reliable and compromised models while explanation fidelity and deployment accuracy differ, so detection requires access to the decision mechanism. They also show that the competence envelope is non-empty and star-shaped under stated validity, continuity, and monotonicity conditions.
- Prediction–explanation separation: For any β ∈ (0,1) and γ ∈ (0,½), models can match every prediction-side certificate while differing in explanation fidelity and deployment accuracy by at least γ.Theorem 1 gives C(f) = C(f′), F(f) = 1, F(f′) ≤ 1 − β, and acc_{P′}(f) − acc_{P′}(f′) ≥ γ.
- Detection lower bound: Prediction-only tests have zero excess power because reliable and compromised models induce identical prediction-law samples; any test exceeding size must evaluate the sensitivity map.The corollary states sup over prediction-only tests of (power − size) = 0.
- Attack interpretation: A dormant-cue data-poisoning attack is theoretically invisible to accuracy and coverage, rather than an empirical accident.The attack plants a cue that is dormant on the clean support and becomes detectable through the separation construction.
- Numerical verification: 0.0 maximum prediction difference coexists with conformal coverage 0.8820 = 0.8820, accuracy 0.8240 = 0.8240, and calibration error 0.0199 = 0.0199.The controlled logistic construction also reports Kolmogorov–Smirnov distance 0.0, explanation fidelity declining from 1.00 to 0.19, deployment accuracy from 0.824 to 0.409, and attack success 0.98.
- Competence-envelope geometry: Under conformal validity and continuity, K(α,β) is non-empty and contains a neighbourhood of the nominal condition; with ray-wise non-increasing C and F, it is star-shaped about 0.Its radial boundary is ∂K(u) = min(t_C(u), t_F(u)).
Supplementary results: architecture-independence
Across five model classes, the competence envelope shows the same stress patterns rather than depending on a single architecture: coverage degrades under drift and explanation-stability drift S rises as data become scarce. In clinical covariate drift, the strongest boosting methods fall furthest outside the certified region on the most out-of-distribution band.
- Architecture-independence: The envelope is evaluated with five model classes using a 100-dimensional truncated-SVD feature projection.The classes are logistic regression, multilayer perceptron (64–32), gradient-boosted decision trees, XGBoost and LightGBM.
- Architecture-independence: Coverage degrades under drift for all five model classes.This pattern is reported for Telegram conformal coverage at n = 2,400 across evaluation months.
- Architecture-independence: Explanation-stability drift S rises as data become scarce for all five model classes.The comparison uses Telegram S at the January anchor across training size n.
- Architecture-independence: The strongest boosting methods fall furthest outside the certified region on the most out-of-distribution clinical covariate-drift band.This result concerns breast-cancer conformal coverage at n = 250.
Supplementary results: foundation models and cross-dataset generality
Supplementary experiments show that the competence envelope generalizes across multilingual foundation-model embeddings and eight public benchmarks. Explanation-stability degradation under scarcity is consistent, while prediction degradation tracks induced covariate shift.
- Foundation-model backbones: Three multilingual transformer backbones reproduced the two-certificate signature on both real corpora.Conformal coverage degraded monotonically under temporal drift, while explanation-stability drift fell with abundance.
- Foundation-model backbones: 0.93→0.86 conformal coverage was observed for DistilBERT-multilingual and multilingual MiniLM on Telegram from January to May.The Telegram evaluation used n=2,400.
- Foundation-model backbones: 0.92→0.85 conformal coverage was observed for XLM-R base on Telegram from January to May.The Telegram evaluation used n=2,400.
- Cross-dataset generality: Explanation-stability degradation under scarcity occurred on all eight public benchmarks using DistilBERT-multilingual embeddings.The envelope used covariate drift induced by principal-component banding.
- Cross-dataset generality: Prediction-certificate degradation tracked induced covariate shift, strongly on heterogeneous multi-topic sets and negligibly on homogeneous sentiment sets.The heterogeneous sets were DBpedia, Amazon, IMDB, and SST-2.
Supplementary Note: reproducibility
The supplementary analyses are reproducible because archived code deterministically generates all reported values from fixed seeds. They rely on public datasets and include explicit faithfulness and conformal-coverage checks.
- Reproducibility: All values in Tables S1–S5 and Figs. 1–2 are generated deterministically from fixed seeds by archived code.The reproducibility pipeline therefore fixes the computational randomness underlying the reported results.
- Data availability: The pipeline uses no proprietary data, drawing on public Telegram and Reddit posts and the public Wisconsin Diagnostic Breast Cancer dataset.The Reddit climate corpus is deposited at doi:10.7910/DVN/NL06IX, while the clinical dataset is distributed with scikit-learn.
- Validation checks: The faithfulness check achieves a linear Shapley efficiency residual ≤ 2×10⁻¹⁵, alongside a conformal coverage target of 0.90.Both checks are specified as part of the reproducibility pipeline.