Source-linked AI summary

Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance

Samar Ansari

arXiv:2609.10105v1cs.CYcs.AIcs.CR

TL;DR

Training-centered compute governance does not fully capture capability that emerges during inference, scaffolding, and local deployment. The paper builds and tests a twenty-mechanism taxonomy using vendor evidence, adversary roles, and governance scenarios. Fifteen mechanisms have production commercial substrates, but coverage is uneven: no mechanism is adequate against a high-capability state-level deployer, and enforcement mechanisms can fail against fine-tuning.

  • Problem

    Training thresholds treat the trained model as the regulatory unit even as capability increasingly shifts to inference scaling, agentic scaffolding, and consumer-hardware deployment.

  • Method

    The paper rates twenty monitoring, verification, and enforcement mechanisms on a four-point scale, using commercial-provider evidence and testing them across adversary roles and governance scenarios.

  • Results

    Fifteen mechanisms have commercial substrates deployable at production scale, but no mechanism is adequate against a high-capability state-level deployer.

  • Takeaways & Limitations

    Inference-time governance is technically more mature than existing regulatory architecture suggests, but deployment readiness and adversary coverage are distinct.

  • Takeaways & Limitations

    The evidence base covers four major commercial providers and does not capture open-weight or self-hosted deployment, where platform-mediated mechanisms lose force.

Abstract

from arXiv · show

Compute governance today is a governance of training: the thresholds, reporting requirements, and frontier-AI regimes now in force attach to training compute and treat the trained model as the regulatory unit. That picture is incomplete: capability increasingly migrates to the deployment stage through inference-time scaling, agentic scaffolding, and compression onto consumer hardware. This paper asks which mechanisms are available once the regulatory object shifts from the training run to the inference call. We develop a feasibility taxonomy of twenty inference-time mechanisms across monitoring, verification, and enforcement, each rated on a four-point readiness scale against a documented four-vendor evidence base. We then stress the taxonomy against a two-dimensional adversary model (three capability tiers crossed with four adversary roles) and map each mechanism to four governance scenarios (domestic regulation, bilateral or multilateral coordination, industry self-regulation, and compute-marketplace governance). Fifteen of the twenty mechanisms have commercial technical substrates in production today, although governance-grade assurance and adversarial robustness vary substantially. The adversary analysis shows that this readiness holds only against a cooperative deployer and a low-to-medium-capability user: no mechanism rates adequate against a high-capability state-level deployer, and fine-tuning removes the model-internal components of the enforcement cluster, although platform-external controls can persist. A substitution analysis connects the taxonomy to a companion hardware paper as a conditional substitution principle describing when inference-stage and hardware-stage mechanisms provide comparable regulatory coverage under stated conditions. A second-rater reliability check on a random subset of the readiness ratings returned a quadratic-weighted Cohen's kappa of 0.74.

1 Introduction

Existing compute governance regulates training runs, but inference-time scaling, agentic scaffolding, and consumer deployment shift capability beyond that unit. The paper responds with a readiness-rated taxonomy that tests mechanisms against separated adversaries and governance settings.

  • The gap: Training-compute thresholds and reporting regimes treat the trained model as the regulatory unit and attach obligations when training ends.This approach was designed for a period when training compute was the dominant capability determinant.
  • The gap: Inference scaling, agentic scaffolding, and consumer-hardware deployment move capability into deployment and weaken training compute as a complete governance proxy.Agentic systems can derive capability from tool-using loops, memory, and external resources as well as from the base model.
  • The gap: The literature identifies the gap but has not assembled technical proposals for verifiable inference, runtime control, and monitoring into a governance framework.The paper positions itself between computer-science proposals that treat governance as code and policy work that treats technical infrastructure as opaque.
  • Approach: The paper classifies twenty inference-time mechanisms across monitoring, verification, and enforcement using a four-point technology-readiness scale.Mechanisms are defined operationally, grounded in academic and vendor sources, and described at their deployment layer.
  • Approach: The adversary model crosses three capability tiers with four roles, distinguishing developers, deployers or integrators, end users, and fine-tuners or scaffold-builders.This role separation captures threats that a one-dimensional capability scale can conflate.
  • Approach: Each mechanism is mapped to domestic regulation, bilateral or multilateral coordination, industry self-regulation, and marketplace governance, then connected to hardware governance through substitution analysis.The convergence analysis states conditions under which inference-stage mechanisms must replace training-time mechanisms.

2 The Inference-Time Compute Landscape

Existing compute-governance instruments focus on training thresholds and largely leave inference unregulated. The landscape therefore points toward deployment-stage, capability-based, and intermediary-focused governance while exposing privacy, open deployment, and operational gaps.

  • Existing instruments: Existing compute-governance instruments assume capability accrues during capital-intensive, geographically concentrated training and focus regulatory attention on the developer.The survey distinguishes threshold regimes, deployment-stage instruments, and newer inference-aware proposals.
  • Existing instruments: Every surveyed threshold regime triggers obligations on training compute, while none attaches an obligation to inference.The threshold-as-trigger pattern is established across the EU, United States, and emerging UK frontier-AI regulation.
  • Threshold limitations: Training thresholds can be bypassed through fine-tuning, component reuse, model expansion, or above-optimal inference-time scaling.Above-optimal inference scaling is distinct because it requires no additional training compute.
  • Threshold limitations: The literature treats thresholds as partial instruments that require capability evaluations and post-market monitoring, while inference-side enforcement remains open.The EU AI Act’s model-centric approach also leaves downstream deployment accountability substantially unresolved.
  • Deployment-stage instruments: Compute-provider intermediation can transfer KYC and record-keeping from training to inference, but inference creates high-volume, fine-grained monitoring and privacy challenges.Downstream fine-tuners, modifiers, and scaffold-builders are also identified as essential governance nodes.
  • Inference-aware proposals: Inference-aware proposals shift attention toward deployed-system capability profiles through evaluation-driven, multidimensional, risk-tiered, or layered frameworks.These proposals share the recognition that capability and compute decouple at inference, though they differ operationally.

3 Taxonomy of Inference-Time Governance Mechanisms

The taxonomy evaluates twenty inference-time governance mechanisms using a four-vendor evidence base and a four-point readiness scale. It spans monitoring, verification, and enforcement, while distinguishing production availability from assurance and adversarial robustness.

  • Evidence base: The evidence base covers Anthropic, OpenAI, Google Vertex AI, and AWS Bedrock, with the first three curated to row level and AWS surveyed at a broader tier.Vendor documentation snapshots are dated May 22, 2026, and AWS coverage counts are lower bounds in the taxonomy.
  • Monitoring: Per-query accounting is currently deployable because all four vendors expose production-scale usage or rate-limit telemetry.The reported substrates include token, reasoning-token, and rate-limit metrics.
  • Monitoring: Per-request expert-activation telemetry is near-term: MoE inference generates the substrate internally, but no vendor exposes it at the user-facing API.Deployment would require caller access and aggregation infrastructure for regulatory use.
  • Monitoring and verification: Caching and regional infrastructure controls are currently deployable, with four-vendor productisation, documented cache accounting, quotas, logs, or compliance features.The cited mechanisms include concrete cache TTLs, per-token billing, regional accounting, and documented certifications.
  • Verification: Energy monitoring and zkML remain less mature: energy monitoring is near-term, while zkML requires R&D because frontier-scale proof generation remains operationally costly.The energy-monitoring gap concerns user-facing product integration; zkML has research demonstrations but no vendor productisation.
  • Enforcement: The enforcement cluster is currently deployable against users and cooperative deployers, but fine-tuning can defeat model-internal defences; external controls may persist.The taxonomy distinguishes production readiness from adversary-scoped assurance, and identifies architectural diversity across runtime enforcement and sandboxing.

4 Feasibility Constraints and Adversarial Considerations

Inference-time governance faces substantial feasibility limits: compute-signature detection erodes with efficiency gains, decentralised deployment weakens platform controls, and adversarial coverage collapses for high-capability deployers and fine-tuners.

  • Substitutability and compute signatures: Inference mechanisms identified by compute signatures are not durably deployable because algorithmic efficiency can reduce the compute signature of high-capability inference.This limitation applies to parts of M1 and M4 and the near-term M2.
  • Distributed and decentralised inference: Distributed inference across untrusted or cross-jurisdictional nodes weakens attestation, monitoring, licensing, and jurisdiction-bounded enforcement mechanisms.The challenge is greatest when each node holds only a fragment of the inference.
  • Distributed and decentralised inference: Fully unfederated open-weight inference on consumer machines has no shared monitoring substrate, shifting fallback governance to hardware-layer accelerator controls.This marks a handoff from inference-time governance to the companion hardware-governance framework.
  • Readiness and scope: Fifteen mechanisms have commercially available production substrates, but this headline describes the regulatory toolbox rather than complete governance assurance or adversarial robustness.Readiness ratings concern substrate availability; assurance and robustness vary separately.
  • Adversarial coverage: The user column falls from thirteen adequate mechanisms at low capability to six at medium capability and three at high capability.The pattern supports user-facing safeguards for casual or moderately sophisticated misuse but provides weaker assurance against state-level users.
  • Adversarial coverage: Seventeen of twenty mechanisms are inadequate against the high-capability state-level deployer, while V1, V4, and E4 retain only partial coverage.No mechanism rates adequate in C3.R2; V4 is the strongest of the three partial mechanisms because its cryptographic signing is less dependent on platform cooperation.
  • Adversarial coverage: Fine-tuning defeats model-internal and model-mediated enforcement at medium and high capability, while platform-external controls can persist.E2 is inadequate at C2.R4 and C3.R4; E5, E6, and E7 fall from partial to inadequate at C3.R4.

5 Convergence between Hardware-Level and Inference-Time Governance

The convergence analysis maps hardware-level and inference-time mechanisms across overlapping but distinct regulatory scopes. It identifies parallel, layered, and architecturally divergent substitutions, culminating in a conditional principle for when stage-specific mechanisms provide comparable coverage.

  • Structural mapping: The two taxonomies cover different regulatory objects but overlap at the cloud-platform layer, where hyperscalers can mediate both hardware and inference governance.Hardware governance targets chips, racks, and infrastructure providers; inference governance targets requests, deployed systems, and marketplaces.
  • Structural mapping: Twelve of twenty inference-time mechanisms have a direct or near-direct hardware analogue, while eight do not.The mapping is tabulated in Table 5.
  • Substitution patterns: Twelve mechanisms substitute across domains: nine are parallel substitutions and three are layered substitutions built on hardware-derived evidence.Seven mechanisms have no hardware analogue, while V2 is architecturally divergent.
  • Conditional substitution principle: Parallel mechanisms provide comparable coverage at either lifecycle stage when the substitution channel is bounded, making intervention stage a policy choice under the stated conditions.This equivalence applies only when the regulated activity occurs predominantly at the selected stage and substitution remains bounded.
  • Conditional substitution principle: Layered mechanisms become necessary when substitution opens because they depend on hardware-level evidence, while architecturally divergent mechanisms require independent deployment at both stages.The layered set is M4, V3, and E3; the divergent inference-side case is V2.
  • Conditional substitution principle: Inference-only mechanisms remain required regardless of substitution dynamics because the principle does not provide hardware analogues for them.The seven listed mechanisms are M3, V4, V6, V7, E2, E5, and E6.

6 Mapping Mechanisms to Governance Scenarios

The scenario mapping shows that governance mechanisms change role according to which actor holds authority. No single scenario is comprehensive: domestic and marketplace systems provide runtime enforcement, multilateral coordination provides cross-border verification, and self-regulation contributes reasoning transparency.

  • Cross-scenario pattern: Six mechanisms are levers in no scenario, while E3 is the only mechanism serving as a lever in three scenarios; none is a lever in all four.The four scenarios are domestic regulation, multilateral coordination, industry self-regulation, and marketplace governance.
  • Domestic regulation: The domestic scenario is mechanism-broad, with sixteen actionable mechanisms and no structural blocks, but its controls remain territorially bounded.Every domestic lever binds cooperative in-territory deployers, not self-hosted state-level deployers or foreign providers outside the jurisdiction.
  • Bilateral or multilateral coordination: Multilateral coordination makes verification mechanisms primary levers because shared certification authorities give cross-border attestations practical value.Distributed-inference monitoring is uniquely a multilateral lever, while jurisdiction-bounded inference operates in three scenarios.
  • Lessons from existing regimes: Existing monitoring regimes contribute structural patterns rather than one-to-one templates: financial regulation informs intermediated KYC, pharmacovigilance informs capability monitoring, and aviation informs incident reporting.The analogues are complementary and none maps directly onto inference governance.
  • Industry self-regulation: Self-regulation is the weakest scenario, with one lever, twelve supporting ratings, four irrelevant ratings, and three structural blocks.Its voluntary commitments formalise productised mechanisms but cannot create the independent authority required for credible cryptographic attestation.

7 Discussion and Implications

The discussion finds that inference-time governance is commercially deployable but unevenly capable: readiness is broad against ordinary cooperative threats and thin against sophisticated or locally deployed adversaries. The paper therefore emphasizes verification productisation, hardware coupling, scenario combination, and explicit scope limits.

  • Readiness gap: Fifteen of twenty mechanisms have commercially available production substrates, while two are near-term, two require research, and one is speculative.The production base comes from multiple major commercial vendors and supports billing, abuse prevention, and capacity allocation.
  • Readiness gap: The deployable substrate is concentrated against cooperative deployers and low-to-moderate-capability users, while no mechanism is adequate against a state-level deployer.At higher capability, fine-tuning also disables model-internal enforcement components, although platform-external controls can persist.
  • Readiness gap: The paper interprets optimistic deployment results and pessimistic adversary-coverage results as simultaneous properties of the same evidence matrix.The taxonomy makes the conditions for each reading explicit rather than collapsing them into one readiness claim.
  • Research priorities: The first research priority is productising the verification cluster, especially exposing existing infrastructure-level attestation substrates through inference APIs.V1 is near-term because confidential-computing substrates exist but are not yet exposed at the inference API.
  • Research priorities: As training-inference substitution becomes more efficient, layered mechanisms M4, V3, and E3 gain regulatory weight because hardware becomes a harder-to-evade intervention point.This couples the inference-time research agenda to hardware-level governance.
  • Policy implications: A comprehensive regime should combine governance scenarios rather than choose among them, pairing runtime enforcement, cross-border verification, and reasoning transparency.No mechanism is a primary lever in all four scenarios, and only E3 is a lever in three.
  • Limitations: The evidence base covers four frontier commercial vendors but excludes open-weight and self-hosted systems, leaving productisation findings silent beyond the commercial frontier.This boundary is especially important because marketplace and self-regulation scenarios cannot reach that surface.
  • Limitations: The reliability estimate is based on seven of twenty mechanisms and is fragile: quadratic-weighted Cohen’s kappa is 0.74, while a single rating change can shift it by roughly 0.13.Exact agreement occurred on five mechanisms, with one-step disagreements on V4 and V7.

B Appendix B: Inter-Rater Reliability Methodology

The reliability check tested whether readiness ratings were reproducible across seven randomly selected mechanisms using an independent technically expert rater and a sealed, identical rating protocol.

  • Sampling: Seven mechanisms were randomly selected before either rater’s ratings were examined, spanning monitoring, verification, and enforcement.The sample was drawn from the final mechanism list.
  • Rater selection: The second rater was an engineer with reinforcement-learning and IoT experience who had not participated in the paper.The out-of-field technical expertise was intended to test whether the readiness signal was legible without the author’s framing.
  • Protocol: The author’s ratings were sealed before the second rater received a self-contained brief without the paper or the author’s ratings.The separation was designed to prevent knowledge of the original ratings from influencing the re-rating.
  • Protocol: Both raters used the same four-point scale and conservative rule, rating mechanisms against the relevant adversary rather than only a cooperative actor.Adjacent uncertainty was resolved by choosing the higher, more conservative rating number.

B.4 Agreement statistic

The seven-mechanism reliability check produced substantial agreement, but the interpretation depends on the prespecified weighting choice and is fragile at this sample size.

  • Five of seven mechanisms received identical ratings; the two disagreements were one-point differences, with the second rater more conservative in both.
  • 0.74 quadratic-weighted Cohen’s kappa indicated substantial agreement under the Landis and Koch interpretation.Quadratic weighting treats adjacent disagreements on the ordered four-point scale as less serious than larger disagreements.
  • The substantial-versus-moderate interpretation depends on weighting: unweighted kappa was 0.53 and linear-weighted kappa was 0.63.
  • The disagreement-resolution procedure applied only to disagreements greater than one point, so the two one-point differences did not revise the primary ratings.

C.1 V4: tool-call cryptographic signing (author 1, second rater 2)

The V4 rating disagreement concerns whether production availability of the core signing primitive is sufficient for a currently deployable end-to-end provenance mechanism.

  • The author rated V4 currently deployable because two of four vendors productise the signing primitive at production scale.The second rater instead rated it near-term because the full provenance chain was not yet demonstrated in production.
  • The disagreement is a threshold choice between two-vendor availability of the core primitive and production readiness of the complete provenance chain.
  • The author retained the deployable rating because the signing primitive is already available at frontier scale and underlies the provenance chain, while recording the conservative alternative.

C.2 V7: chain-of-thought monitorability (author 1, second rater 2)

V7 illustrates a distinction between deploying reasoning-trace exposure and establishing that those traces faithfully represent model computation, with the latter limiting adversarial assurance.

  • The author rated V7 currently deployable because three of four vendors expose reasoning traces, summaries, or signed thinking through an API or console.
  • The adversary analysis reconciles the rating by treating V7 as deployable-as-primitive but weak-as-defence until trace faithfulness is resolved.
  • The second rater identified V7 as the hardest mechanism to rate because exposing reasoning traces may not establish that they faithfully represent computation.
  • The evidence base is documented across AWS Bedrock, Anthropic, OpenAI, and Google Vertex AI, with Azure OpenAI excluded as structurally overlapping.
  • Vendor evidence was captured in dated snapshots, and documented absences were retained as evidence for adversary-model implications.

D.1 M1: Per-query usage and token accounting

Per-query usage and token-accounting mechanisms draw on vendor APIs, rate limits, token-counting utilities, and billing fields, but their granularity and semantics vary across providers and interfaces.

  • Token counts apply to text inputs only in the cited limitation, leaving image and other modalities outside that accounting scope.
  • Usage controls include rate limits, token-counting utilities, per-call metadata, usage reporting, and per-organization tiers.The evidence includes token and request dimensions, pre-flight counting, and provider-specific usage endpoints.
  • Accounting limitations include end-of-message rather than per-chunk thinking-token reporting, API-version-dependent fields, opaque tier auto-graduation, and model-specific tokenization conventions.
  • Per-call and billing records distinguish input, output, cached, reasoning, and tool-execution usage in provider-specific ways.Examples include reasoning-token fields, cache-aware accounting, and separate code-execution inputs and outputs.
  • $200,000/month is the documented ceiling for OpenAI Tier 5 usage limits.
  • Reasoning-token accounting does not expose the underlying reasoning content, while OpenAI exposes reasoning summaries rather than raw chain-of-thought.

D.2 M3: KV-cache and reasoning-token accounting

Inference-time accounting exposes controls for cached context, reasoning depth, and token use, but their granularity, documentation, and model support vary across vendors.

  • KV-cache accounting: Cache breakpoints, TTLs, utilization diagnostics, and context-management primitives provide deployer-visible controls for repeated or long conversations.Anthropic documents cache diagnostics, breakpoint limits, TTLs, and per-cache controls; comparable context-cache primitives are described for other routes.
  • Reasoning-token accounting: Manual reasoning budgets are inconsistently available: the manual budget mode is deprecated on some Anthropic models and removed on others.The cited documentation distinguishes Sonnet 4.6 support from removal on Opus 4.7/4.8/Fable/Mythos.
  • Reasoning-token accounting: Reasoning depth can be exposed through adaptive effort parameters, thinking levels, or token budgets, but these controls are often coarse or model-dependent.Documented examples include Anthropic effort and task budgets, Gemini thinking_level, and OpenAI reasoning_effort.
  • Reasoning-token accounting: OpenAI reasoning-effort selection provides guidance rather than enforcement, while reasoning models remain slower and more expensive.The documented trade-off is explicit, but customer-side selection does not constitute a hard server-enforced token ceiling.
  • Context and latency controls: Server-side compaction and predicted outputs can reduce context or latency, but compaction is opaque and predicted outputs help only when predictions substantially match.Compaction preserves some history while trading older content for newer content; predicted outputs otherwise incur billing without benefit.

D.3 M4: Distributed inference monitoring across data centres

Distributed inference monitoring is partly supported through jurisdiction controls, regional processing, and residency features, but telemetry, quota, and routing limitations weaken centralized visibility.

  • Jurisdiction and routing: Anthropic’s inference_geo parameter selects US or global inference, while rate limits remain shared across geographic values.Cloud-platform variants may expose different regional behavior, and documented region availability varies by platform.
  • Jurisdiction and routing: Google documents region selection and residency-related processing controls, but some operations may use multi-region storage or processing.The cited material distinguishes processing-region selection from a guarantee that data remains within one region.
  • Jurisdiction and routing: OpenAI lacks native multi-region inference monitoring on the standard platform, including per-region telemetry and inference-geo-separated rate limits.The closest documented analogue is project-level data-residency control rather than direct monitoring of individual inference locations.

D.4 M5: Authenticated account identity and request tracing

Account identity, auditability, confidential computing, and authenticated tool-call mechanisms provide partial request-tracing substrates, but access, retention, and verification limits constrain assurance.

  • Auditability: Audit logging and prompt-retention controls exist across cloud and API platforms, but logging is often opt-in, variably retained, or unavailable on some deployment routes.Examples include project-level Cloud Audit Logs, standard-tier prompt retention, and route-specific availability limits.
  • Auditability: OpenAI provides abuse-monitoring logs and administrative APIs, but retention depth and granularity are not fully documented publicly.Administrative operations require a separate sensitive credential, and retention duration is not explicitly bounded in the cited policy material.
  • Confidentiality: Confidential-computing products offer hardware-backed encryption-in-use, including documented Gemini inference availability, but serving coverage is limited by product and location.The evidence distinguishes Confidential VM and related products from unavailable or narrower user-facing serving configurations.
  • Authenticated outputs: Tool-call and reasoning signatures can authenticate model-generated artifacts, yet their opaque formats, incomplete public key documentation, and non-cryptographic schema enforcement limit independent verification.Anthropic documents signatures on thinking blocks, while strict tool use enforces declared schemas without cryptographic validation.

D.7 V5: Capability-evaluation reporting

Capability-evaluation reporting is supported by programmatic evals, graders, retrieval citations, reasoning traces, access controls, and safety classifiers, but portability, provenance, transparency, and robustness remain limited.

  • Evaluation infrastructure: OpenAI’s Evals API supports reusable graders and data sources, but the platform is deprecated, existing eval IDs are not portable, and results are model-graded data.The cited documentation gives a read-only and shutdown timeline alongside portability and interpretation limits.
  • Grounding and provenance: Retrieval and web-search tools return source URLs or snippets, but they lack cryptographic corpus provenance, tamper evidence, or guaranteed citation accuracy.These limitations recur across first-party RAG, search grounding, web search, and vector-store implementations.
  • Reasoning observability: Reasoning observability varies sharply: Anthropic can expose thinking content, whereas OpenAI exposes summaries and token counts rather than raw chains of thought.Anthropic also offers display modes that may summarize or omit raw traces; OpenAI’s caller-side monitorability is structurally limited.
  • Access and enforcement: Usage tiers, IAM, project permissions, spend limits, and model gating provide access-control substrates, but automatic tier advancement and some ceilings are opaque or constrained.The evidence includes fixed throughput, per-resource permissions, organization tiers, and hard monthly ceilings alongside opaque graduation rules.
  • Safety enforcement: Safety classifiers, refusals, and acceptable-use policies support enforcement, but classifier accuracy, refusal taxonomies, appeals, and opt-out behavior are incompletely characterized.The cited materials describe classifier-based refusals and computer-use screening while noting undisclosed accuracy, policy gaps, or hostile-deployer opt-out paths.

D.13 E4: Rate limiting as governance primitive

Rate limiting is a commercially deployed inference-time governance primitive spanning requests, tokens, queues, batches, service tiers, and acceleration modes. Its governance value is strongest for cooperative use, while silent downgrades, contention failures, and structural rather than semantic controls limit assurance.

  • Governance boundary: Priority and fast modes provide accelerated service but remain subject to tier or ramp-rate constraints and may fail during contention.OpenAI documents a ramp-rate limit and resource-unavailable failures for flex processing, while Anthropic documents dedicated fast-mode limits.
  • Commercial substrates: Batch and asynchronous endpoints add queueing controls, including separate queue limits, processing windows, and maximum requests per batch.Anthropic and OpenAI document distinct batch queues and per-batch limits; OpenAI’s batch endpoint uses a 24-hour completion window.
  • Commercial substrates: Tiered service and quota systems let providers govern inference throughput through named, globally bounded metrics and per-request service choices.Anthropic documents tier-based limits, while Google documents standard, priority, flex, and reserved throughput options.
  • Commercial substrates: Vendors expose multiple rate-limit dimensions, including requests, tokens, queue depth, batch size, and processing tiers.OpenAI documents RPM, RPD, TPM, and TPD limits, while batch APIs add queued-token and per-batch constraints.
  • Governance boundary: Rate limiting does not guarantee stable enforcement because silent downgrades, token-bucket behavior, and resource contention can alter user-visible service.Priority processing may silently downgrade to standard, and nominal limits can still produce 429 responses under token-bucket enforcement.

D.15 E6: Tool-use authorization and sandboxing

Tool authorization and sandboxing provide deployment-side controls through structured tool surfaces, isolated execution, quotas, logging, and approvals. Their enforcement strength depends heavily on caller or customer implementation, with several controls remaining advisory, opt-in, or incompletely documented.

  • Enforcement boundary: Authorization and isolation are often application- or customer-side rather than provider-enforced.OpenAI states that tool selection and sandbox details remain caller responsibilities, while Anthropic provides tool schemas without implementing execution or sandboxing.
  • Sandboxing: Sandboxing can isolate execution environments and impose regional quotas on execution, writes, entities, and agent-runtime operations.Google documents per-region sandbox quotas including 1000 RPM execution and 500 RPM writes.
  • Tool surfaces: Tool-use systems expose structured action surfaces for computer control, code execution, shell access, file editing, and MCP integrations.These interfaces define available tools or actions, but execution and authorization responsibilities vary by platform.
  • Governance boundary: Several controls have limited assurance because capabilities vary by model, sandbox details are unpublished, hosted MCP servers lack provider attestation, and enforcement criteria or appeals are incompletely documented.These boundaries constrain portability and governance-grade assurance across vendors and deployments.
  • Monitoring: Logging can support post-hoc audit, but documented implementations may make it opt-in and place storage or configuration responsibilities on the deployer.Vertex AI provides configurable request-response logging containing prompts, responses, latency, and status.
Loading 2609.10105v1…