Source-linked AI summary
Compute-Bounded Security Assurance - Coverage, Verification, and Response under Resource Constraints
Jithin VG, Ditto PS
TL;DR
Security-assurance studies need to distinguish repeated success, distinct coverage, accepted evidence, and operational protection under finite resources. The paper integrates coverage theory, evidence checking, scoring, accounting, capacity, response, and evaluation into a framework whose central result is that pairwise correlation does not determine an asymptotic coverage ceiling. Its scope is theoretical and analytic rather than an empirical scaling law.
Problem
Security-assurance progress cannot be characterized reliably by generated volume or compute alone because coverage, correctness, evidence acceptance, cost, capacity, and response timing are distinct quantities.
Method
The paper builds a resource-constrained assurance framework combining latent-success coverage models, fallible checking, proper scoring, complete resource accounting, service-capacity analysis, response latency, and evaluation protocols.
Results
Pairwise outcome correlation determines a variance adjustment, not a universal coverage ceiling; under conditional-iid repetition, the asymptotic ceiling is determined by zero-probability support.
Takeaways & Limitations
Finite-budget comparisons should report correctly adjudicated, distinct, timely outcomes under complete resource constraints rather than extrapolating coverage or protection from correlation or compute alone.
Takeaways & Limitations
The architecture is conceptual: no implementation, calibrated population parameters, hardware comparison, model-performance study, or deployment trial is reported.
Abstract
from arXiv · showhide
Additional inference compute can increase the number of correctly resolved security-assurance tasks, but repeated success, unique coverage, accepted evidence, and operational protection are different quantities. We develop a resource-constrained framework that separates them. For repeated conditionally independent attempts with latent success probability $Θ$, coverage is $C_n = 1 - E[(1-Θ)^n]$, and its limiting value is $1 - P(Θ= 0)$. Positive pairwise outcome correlation does not by itself imply a ceiling below one: we construct two models with the same mean success and pairwise correlation but different limiting coverage. We distinguish this result from the effective sample size used to estimate a mean, and show why finite-budget observations cannot generally identify an asymptotic support ceiling. We then connect coverage to fallible evidence checking, proper scoring of factual grounding, complete resource accounting, service capacity, and a response model that includes mitigation delay. A conceptual defensive architecture separates evidence analysis, adjudication, and operational authority. An evaluation protocol specifies held-out tasks, paired comparisons, negative cases, and uncertainty reporting. The contribution is a consistent theoretical synthesis and a set of counterexamples to invalid extrapolations, rather than an empirical scaling law. All numerical illustrations are analytic; no model-parity result, hardware benchmark, or general attacker-defender equilibrium is claimed.
1 Introduction
The paper frames security assurance as a resource-constrained process of producing and validating evidence about explicit obligations, rather than equating more generated text or compute with security progress. It integrates coverage, verification, cost, and operational distinctions while using counterexamples to block unjustified extrapolation.
- Motivation: More generated reports can coexist with fewer resolved obligations, greater reviewer burden, or missed deployment deadlines.The paper therefore treats output volume as an ambiguous measure of security progress.
- Motivation: Additional inference can improve task coverage in particular settings, but existing evidence does not establish a universal FLOPs-to-security law or correlation-based coverage limit.Repeated-sampling studies and security demonstrations motivate resource allocation without proving universal scaling relationships.
- Scope: Security assurance concerns explicit obligations and adjudicated outcomes, not autonomous exploitation procedures.The framework analyzes repeated resolution attempts abstractly and treats harmful-event processes only as response-timing models.
- Contributions: The framework separates unverified obligations, actual violations, proved satisfaction, accepted reports, service capacity, and response latency.Its contributions integrate coverage modeling, error-aware verification, complete resource accounting, constrained capacity, and response timing.
- Contribution boundary: The contribution is an integrated assurance model with explicit counterexamples, not new probability theory, empirical performance results, or transferred component speedups.Related systems and mathematical tools are treated as foundations or component results rather than evidence of end-to-end security improvement.
- Scope of inference: Composition bounds and coverage measures do not by themselves establish lower bounds, monotonicity, violation density, or complexity laws.The paper emphasizes that upper bounds on reachable states and counting identities require additional growth or modeling assumptions.
4 Sequential assessment and evidence quality
The paper treats sequential assessment as a pipeline in which uncertain observations, probabilistic predictions, participation, correctness, and checker acceptance must remain distinct. It recommends explicit stochastic modeling, proper scoring, and error-aware accounting rather than collapsing these quantities into one capability measure.
- Sequential assessment: A partially observable model represents transitions, observations, rewards, initial beliefs, and resource or assessor state, while retaining stochastic tools and incomplete evidence.The policy remains separate from the environment, and a history-state MDP is an alternative when history predicts future observations.
- Sequential assessment: Completion means a correct adjudicated conclusion about an obligation, whereas intermediate activity measures need not preserve the terminal objective.The defensive scope includes reviewing authorized evidence, checking configuration invariants, and assessing documented remediation.
- Acceptance and correctness: Accepted evidence is not equivalent to correct evidence: P(A = 1) = πv + (1 − π)f and P(Z = 1 | A = 1) = πv / [πv + (1 − π)f].Here π is correctness prevalence, v is acceptance conditional on correctness, and f is false acceptance conditional on incorrectness.
- Acceptance and correctness: With π = 0.01, v = 0.90, and f = 0.01, the positive predictive value is 0.476190.The example shows why high sensitivity alone does not ensure reliable accepted reports; base rates and conditional errors must be measured on the evaluated population.
- Error accounting: Checker error guarantees require conditional bounds over adaptive histories, and repeating checkers with shared failure modes does not justify multiplying their error probabilities.A repeated-checking union bound requires no independence, but its premise is stronger than a favorable static-validation false-acceptance rate.
- Grounding and scoring: Grounding quality should use proper predictive scoring with declared preprocessing and support conventions, while separately scored facts should not be treated as joint likelihoods without a joint model.The paper also warns that score differences and related rates may have different units, signs, and interpretations.
- Participation and correctness: Participation changes do not establish improved conditional correctness, so refusal, missing information, competence, and acceptance should not be collapsed into one scalar capability measure.The exact chain rule keeps participation and correctness as distinct factors.
5 Coverage under repeated assessment
Repeated assessment coverage depends on the distribution of per-outcome success probabilities, not merely mean success or pairwise correlation. The framework distinguishes eventual coverage from practical budget limits and shows why finite observations cannot generally identify an asymptotic ceiling.
- Define coverage: Coverage is defined over a finite evaluation universe of distinct weighted outcomes, with each outcome counted when at least one attempt correctly resolves it to the declared evidence standard.The measure differs from finding any issue in a system or counting accepted reports, and requires a known evaluation universe or sampling frame.
- Coverage model: Under conditionally iid attempts with per-outcome success probability p_j, repeated coverage aggregates 1 − (1 − p_j)^n across weighted outcomes.A latent success probability Θ gives the equivalent form C_n = 1 − E[(1 − Θ)^n] for fixed-procedure repetition.
- Coverage limits: The limiting coverage equals one minus the probability of zero latent success, while positive success probabilities can still create severe practical budget or deadline barriers.Thus a strict asymptotic ceiling and an operationally inaccessible region are distinct claims.
- Heterogeneity: Replacing heterogeneous difficulty by its mean is optimistic for repeated coverage, and the common-rate curve is exact only when all weighted outcomes share the same success rate.The parameter λ = −log(1 − p) is not generally the one-attempt coverage fraction; substituting p is only a small-p approximation.
- Correlation is insufficient: Pairwise correlation is a variance identity for estimating a mean, not a determinant of the all-failure probability or limiting coverage.For any q, c ∈(0, 1), infinitely extendible conditionally iid models can share mean and pairwise correlation while having coverage limits one and strictly below one.
- Correlation is insufficient: For q = c = 0.1, two models have limits 1 and 10/19 ≈0.526316, while the effective-sample-size substitution predicts 1 −0.910 ≈0.651322.The models share mean success and pairwise correlation, yet their coverage curves and limits differ; diminishing increments occur with and without a strict ceiling.
- Finite-budget inference: Finite-budget observations can be made arbitrarily close under models whose asymptotic coverage differs substantially, so exact zero-probability mass requires additional structural or parametric assumptions.Moving an atom at zero to ε raises limiting coverage by π0 while remaining difficult to distinguish at any fixed observation budget.
- Beyond iid repetition: Independence is not necessary for eventual success: failure-conditioned success probabilities bounded below by ε, or having divergent sum, suffice for coverage to tend to one.This result concerns one outcome and histories of positive probability, whereas adaptive reassessment that changes information or procedure requires a different model.
6 Substitution is conditional on cost, support, and time
Substitution claims depend jointly on success support, complete costs, and deadline capacity, not on success probabilities alone.
- Success and support: 29 attempts at p = 0.1 and 149 at p = 0.02 achieve s = 0.95, a ratio that does not determine compute cost.The attempt ratio is 149/29 ≈5.138.
- Complete cost: Cost accounting includes failed attempts, adjudication, shared costs, and per-attempt costs, so a lower-success system can be cheaper or more expensive.Illustrative cost ratios are 0.514 when attempts cost one tenth as much and 8.221 when they cost 1.6 times as much.
- Complete cost: Effective-sample-size discounting does not reduce the physical compute charged for redundant attempts.Economic models must charge raw attempts even when statistical dependence reduces information.
- Success and support: Targets above supported fraction g are impossible, while s = g is approached but not reached at finite n when 0 < p < 1.Heterogeneous tasks require the full coverage model rather than the two-parameter simplification.
- Time constraints: By deadline D, m identical slots with duration t can complete at most m⌊D/t⌋ attempts, before preprocessing, variable durations, or checking bottlenecks.Substitution claims must compare the same task distribution and adjudication standard under explicit budget and deadline constraints.
7 Budgets, service capacity, and delivered outcomes
Delivered assurance must optimize correctly adjudicated outcomes under explicit resource, service-capacity, quality, and deadline constraints rather than nominal compute alone.
- Outcome objectives: A constrained policy maximizes expected timely distinct outcomes subject to a budget and a false-acceptance target.The budget may be imposed pathwise or as an expected-budget constraint, and feasibility includes operational and evidence-access restrictions.
- Outcome objectives: Coverage per FLOP is descriptive efficiency, not a sufficient objective: without quality or service requirements, its ratio can favor arbitrarily little work.Fixed-budget maximization, cost minimization at a required outcome level, and net-benefit maximization are distinct problems.
- Resource accounting: Resource accounting should report FLOPs, elapsed time, dollars, and joules separately because hardware can change time or energy without changing specified arithmetic.Hardware can indirectly change which procedures fit a deadline.
- Allocation limits: At a regular interior optimum with binding budget, the allocation equations provide marginal conditions rather than an algorithm or global-optimality guarantee.Integer breadth, bilinear budgets, nonconcavity, and boundary optima require additional treatment.
- Service capacity: Stable open-pipeline operation requires admitted arrival and service demand to remain below available capacity, with additional conditions for general networks.The capacity inequality counts cases, not successful findings, and fixed bottlenecks are not universal asymptotic bounds.
- Delivered outcomes: Distinct timely value requires joint conditional accounting, because multiplying separately estimated marginal fractions generally gives the wrong result.Weighted or multiple-output cases should use expected distinct timely value per completed case.
- Service capacity: In the synthetic three-stage tandem, 10 cases/s is the throughput supremum and 4 outcomes/s the goodput supremum after stipulated distinctness and adjudication yields.The limiting input rate is not stable in this model, and actual yield can depend on load.
- Systems effects: Cache reuse, phase disaggregation, and bottleneck assumptions must be evaluated with transfer, compatibility, queueing, quality, and tail-latency costs included.A hit ratio alone does not determine throughput improvement, and workload regimes do not impose immutable phase laws.
8 From verified outcomes to defensive protection
Verified outcomes protect systems only through a timed chain from detection to effective mitigation, with explicit risk, causal, and strategic assumptions.
- Response timing: Prevention occurs when detection time plus mitigation time is less than harmful-impact time.The model treats detection and mitigation as separate stages rather than equating detection with prevention.
- Response timing: At α = 0.02 s−1 and γ = 0.10 s−1, prevention is 0.833333 with zero delay and 0.457343 for m = 30 s.As detection becomes infinitely fast with fixed delay, the model tends to e−αm.
- Response timing: The response values are conditional on assumed impact hazards, observability, mitigation effectiveness, and deployment delay, not measurements or adversary-compute bounds.Changing those assumptions changes the result.
- Risk and causality: Risk reduction cannot generally be inferred by multiplying per-attempt success or event hazards by a residual-risk fraction.Removing difficult cases can raise success among survivors while reducing absolute harmful opportunity.
- Risk and causality: Distinct finding counts do not identify a causal risk-reduction rate; evaluation requires common-population potential outcomes and randomized or otherwise justified designs.Simulation results must be labeled as simulations.
- Strategic scope: A defender-first discovery probability λD/(λD+λA) is a race identity, not a strategic equilibrium, and it omits post-discovery mitigation.An equilibrium additionally requires strategy sets, information, payoffs, and mutual best responses.
- Strategic scope: Defensive investment bounds do not transfer universally across decision problems because portfolio choices require explicit loss, cost, and intervention models.Diminishing returns alone do not imply a universal 1/e fraction.
9 A conceptual architecture for defensive assurance
The proposed architecture separates evidence analysis, adjudication, and authorized operational change, with auditability and independent policy boundaries across the pipeline.
- Architecture: The architecture is a design specification for authorized assessment and remediation review, not an implemented or benchmarked system.Its trust boundaries separate information, evidence, and authority.
- Evidence and assessment: Versioned evidence feeds bounded assessment, while observations, hypotheses, and proved claims remain distinct and unresolved conflicts stay explicit.Untrusted repository text, retrieved documents, telemetry, and model output cannot expand assessment authority or scope.
- Adjudication: Adjudication records claim scope, supporting evidence, limitations, and status across proof, bounded observation, probabilistic judgment, or unresolved claim.Duplicate outcomes are reconciled through a declared equivalence relation with retained evidence, and error rates include negative cases.
- Authority and enforcement: Adjudicated findings inform authorized change, but model output does not grant permission to act and protective controls remain active when services fail.Time-critical enforcement relies on independently specified bounded policies.
- Architecture: Figure 3 depicts data or reviewed recommendations moving through evidence, assessment, adjudication, and change processes without arrows conferring authority.Audit and measurement span the stages, while deployment remains subject to independently defined policy.
- Serving and isolation: Serving records must include model and runtime versions, complete resource use, compatibility, queueing, cache identity, and access-control scope.Separate prefill and decoding resources require comparison against a colocated baseline under the same requirements.
- Measurement and change control: Audit links each unique outcome to evidence, adjudication, timestamps, resource charges, and deployment, while configuration changes trigger new evaluation.Improvements are judged by correctly adjudicated outcomes and deadline performance, with deployment impact measured separately.
10 Evaluation and reproducibility
The evaluation protocol emphasizes held-out, paired, and resource-complete comparisons that distinguish correctness, coverage, evidence quality, timeliness, and productivity. It also requires explicit estimands, uncertainty procedures, and sensitivity analyses rather than unsupported extrapolation.
- Task construction and estimands: Evaluation should combine exact-ground-truth synthetic obligations, versioned authorized tasks, and sanitized operational replay or shadow-mode recommendations.These settings assess calibration, duplicate accounting, abstention, specification and configuration reasoning, remediation assessment, timeliness, and reviewer burden, but do not directly estimate protection in an unobserved deployment.
- Task construction and estimands: Before comparisons, declare task populations, versions, outcome equivalence, weights, admissible evidence, budgets, deadlines, and primary estimands.Time-held-out tasks, separated tuning and test projects, contamination documentation, and blinded reviewers improve interpretability.
- Task construction and estimands: Correctly adjudicated obligations, violations, unsupported claims, failed runs, timeouts, unavailable evidence, and abstentions must remain separately visible.For an intention-to-assess endpoint, excluded or unresolved cases remain in the denominator.
- Coverage and extrapolation: Coverage should be estimated directly over observed budget grids, while higher-budget extrapolations require held-out validation and residual reporting.Pairwise correlation, trajectory similarity, root-cause overlap, and new-outcome fractions are distinct estimands; all-success and all-failure strata require special handling.
- Coverage and extrapolation: Multiple response families should be fit where supported, with sensitivity reported because asymptote intervals are conditional on the chosen family.Observed plateaus may reflect censoring, insufficient budgets, task heterogeneity, checker bottlenecks, or changing conditions.
- Resource accounting: Report distinct timely outcomes, error types, grounding, utility, case completion, latency tails, and complete resource use rather than fictitious FLOPs.Accounts should include preprocessing, inference, tools, adjudication, retries, failures, and transfer overhead; human review time should be reported separately.
- Paired comparisons and uncertainty: Noninferiority requires a valid one-sided confidence bound for ∆=PB−PA to exceed −ϵ under prespecified assumptions.Equivalence additionally requires rejecting both directional differences; failure to detect a difference is not equivalence.
- Paired comparisons and uncertainty: Paired outcomes, project-level dependence, preregistered endpoints, and separate cost frontiers are necessary for defensible uncertainty and comparison.Equal-dollar, equal-energy, and equal-deadline comparisons answer different questions and should be reported separately.
11 Limitations and implications
The paper’s coverage claims hold within explicit models, while practical assurance remains limited by finite-budget inference, evidence and queueing assumptions, and the absence of implementation or deployment validation. Its implication is to measure evidence quality, verification, controls, and response according to demonstrated contribution.
- Scope and assumptions: Coverage results are exact only under conditional independence, a latent success distribution, and a fixed evaluation universe.These are modeling assumptions rather than observations about all security work.
- Scope and assumptions: Finite-budget observations cannot generally distinguish inaccessible support from extremely rare success, limiting inference about asymptotic ceilings.Adaptive procedures, open-ended production settings, and unknown coverage denominators further constrain the framework.
- Evidence and response assumptions: Evidence accounting does not solve specification errors or guarantee independent adjudication, while proper scoring requires reliable labels and a defined prediction space.Queueing and response examples also rely on compatible workload assumptions, independent clocks, constant hazards, and effective mitigation.
- Empirical scope: The architecture is conceptual: the paper reports no implementation, calibrated population parameters, hardware comparison, model-performance study, or deployment trial.Formula checks validate arithmetic and displayed examples, not the empirical adequacy of the assumptions.
- Implications: The practical implication is to invest in evidence quality, scoped verification, effective controls, and timely response according to measured contribution.Additional compute is one input, and no universal ranking follows without costs, effectiveness, and an evaluated loss model.
12 Conclusion
Compute-bounded assurance should evaluate distinct, correctly adjudicated, timely outcomes under complete resource constraints rather than equating correlation, repeated success, or productivity with protection. The framework links coverage limits, verification error, service capacity, and mitigation delay while avoiding claims of universal substitution or equilibrium.
- Conclusion: Assurance should be evaluated through correctly adjudicated, distinct, timely outcomes under complete resource constraints.Pairwise correlation adjusts variance, but does not establish a universal coverage ceiling; finite budgets may confuse zero support with extremely rare success.
- Conclusion: Verification error, service capacity, and mitigation delay determine how much assessment activity becomes useful evidence and protection.The framework supports empirical comparisons without claiming inevitable vulnerability, universal model substitution, or a general security equilibrium.
Research transparency
The paper’s numerical illustrations are deterministic evaluations of formulas using synthetic parameters, not empirical assessments of software or operational performance.
- Reproducibility: The numerical illustrations use synthetic parameter choices and report no reproduced vulnerability, live-system assessment, or operational performance data.The companion source package includes calculation and figure-generation scripts.
A Notation
This section defines local notation for resource allocation, analytic checks, and counterexamples used to delimit the paper’s claims. The restricted allocation result relies on separability, concavity, and exponential benefits, while the examples show why broader investment conclusions require additional assumptions.
- Restricted concave allocation model: The allocation model assigns nonnegative budgets x_i to nonoverlapping assurance-benefit classes under a total budget.The benefit function is concave, and the resource budget B uses the same unit as each x_i.
- Restricted concave allocation model: Funded classes satisfy A_iκ_i e^(-κ_i x_i) = η for a common η > 0, with KKT conditions sufficient under the stated assumptions.Positive-benefit coordinates share a common marginal-benefit multiplier.
- Restricted concave allocation model: The multiplier is uniquely determined for B > 0 because the relevant sum decreases continuously from infinity to zero over its positive range.If B = 0, allocations vanish; if all A_i = 0, every feasible allocation is optimal and zero spending is a natural choice.
- Scope and analytic checks: The allocation result is not a bilinear depth–breadth model, and changes involving shared outcomes, deadlines, checker capacity, or uncertain effects require a different optimization model.Analytic scripts supplement the proofs by checking endpoint cases, moments, coverage bounds, and allocation conditions, but do not formally verify the whole manuscript.
- Analytic checks and numerical examples: Bounded concave benefit alone cannot establish a universal 1/e expenditure bound: the counterexample yields x* = 0.45 + log(2)/20 ≈ 0.484657 > 1/e.The example is continuously differentiable, nondecreasing, concave, starts at zero, and is bounded above by one.