Source-linked AI summary

Toward a Threat Actor Profiling Taxonomy for Pre-Release Risk Management of Open-Weight Frontier Models

James Zhang

arXiv:2608.25361v1cs.CY

TL;DR

Threat-actor realism depends on specifying who the potential actor is, while existing profiles are organized around retrospective observed behavior. The paper proposes a six-attribute taxonomy and a common structure for adversary analysis, with particular urgency for open-weight model release decisions.

  • Problem

    Threat-actor realism is actor-relative, but existing profiles are organized retrospectively around observed behavior rather than specified ex ante.

  • Method

    The paper proposes a six-attribute taxonomy in which a threat-actor profile is a filled vector combining specific tiers across all six attributes.

  • Results

    The resulting complete threat-actor profile is notably jagged, with time horizon and prior domain knowledge among its distinguishing attributes.

  • Takeaways & Limitations

    The taxonomy provides a structure for adversary analysis across the field and is particularly urgent for open-weight model developers.

  • Takeaways & Limitations

    The open-weight versus closed distinction is reductive because model access exists on a spectrum and is further complicated by documentation, licensing, and safety mitigation.

Abstract

from arXiv · show

Pre-release risk management for frontier AI misuse risks routinely leaves threat actor assumptions implicit, inconsistently specified, or ungrounded. This capstone argues that explicit adversary characterization should be regarded as a prerequisite for evaluations that are interpretable, comparable, and faithful to the risks they target. We propose a six-attribute taxonomy (covering technical sophistication, prior domain knowledge, organizational capacity, operational infrastructure, financial capacity, and time horizon) with empirically grounded tiers derived from existing terrorism, biosecurity, and cybersecurity literature. The taxonomy is designed to function as research infrastructure: a common language for pre-specifying adversary assumptions before evaluations are conducted, analogous to pre-analysis plans for randomized controlled trials (RCTs) in medicine and economics. Its application is particularly urgent for open-weight model developers, for whom release decisions are irreversible and must anticipate adversarial reasoning.

Introduction

Open-weight models are improving toward frontier capabilities while making misuse controls harder to retain after release. Existing evaluations can identify danger but may not provide sufficient assurance because divergent adversary assumptions make results difficult to interpret and compare.

  • Their dual-use capabilities raise concerns about chemical, biological, radiological, nuclear, and offensive cyber misuse.
  • Open-weight models can be used and modified without authoritative oversight after release, allowing safeguards to be disabled or harmful capabilities introduced.
  • Open-weight releases create pressure for accurate pre-release decisions because they are harder to protect and cannot be recalled.
  • Capability evaluations have a critical blind spot: they can show that an AI system is dangerous but cannot show that it is safe.
  • Existing behavioral frameworks model how threat actors operate retrospectively, whereas ex-ante attribute-based profiling could strengthen risk analysis.
  • The paper proposes making attribute-based adversary profiling a standardized component of pre-release practice alongside safety benchmarking for open-weight developers.
  • A six-attribute taxonomy is intended to help developers reason more precisely about risks and design more targeted, accurate, scoped, and interoperable evaluations.

1.2 Methodology

The paper develops a literature-grounded threat actor taxonomy and explains how it can inform pre-release evaluation design and practice, with particular attention to open-weight models.

  • Research design: The research design begins with a literature review identifying gaps in frontier AI misuse risk management.
  • Research design: It formulates threat actor attributes from cybersecurity and CBRN analysis, ensuring that the attributes are distinct and independently meaningful.
  • Research design: Each attribute receives multiple tiers grounded in established terrorism and cybersecurity theories or, where literature is sparse, observable empirical criteria.
  • Application scope: The taxonomy is framed around open-weight models because publicly downloadable weights cannot be rescinded, monitored, or moderated after release.
  • Application scope: Open-weight releases are especially consequential because developers lose the ability to update filters, revoke access, monitor use, or shut systems down.
  • Evaluation context: Current evaluations include capability testing, threat modeling, expert consultation, and escalation thresholds, but their assumptions and comparisons remain difficult to interpret.
  • Limitations: Comparisons require shared counterfactual baselines that are often difficult to establish, while risk-management commitments remain largely untested under pressure.

2.3 Evaluating Frontier AI Misuse Risks

Frontier misuse evaluations use automated assessments, expert red-teaming, and human uplift trials, each encoding different threat-actor assumptions and design tradeoffs. These methods provide useful evidence but remain sensitive to study design, limited in interactive realism, and often opaque or difficult to compare.

  • Evaluation methods: Misuse evaluations for CBRN and offensive cyber capabilities primarily use automated assessments, expert red-teaming, and human uplift trials.
  • Automated assessments: Automated assessments measure model capabilities reproducibly and at scale through questions, open-ended prompts, and multi-step agentic tasks.
  • Automated assessments: These assessments include LAB-Bench, WMDP, Cybench, and SecureBio VCT, covering bioresearch, security knowledge, cyber tasks, and virology troubleshooting.
  • Automated assessments: Automated evaluations are limited in modeling the interactive, iterative nature of real-world misuse processes.
  • Expert red-teaming: Expert red-teaming simulates adversaries to probe planning, idea selection, and laboratory-assistant capabilities, but results depend on personnel and provided resources.
  • Human uplift trials: Human uplift trials compare model-assisted treatment groups with controls and interpret performance differences as evidence of the model’s contribution to real-world risk.
  • Human uplift trials: Approximately four times more accurate performance was observed for LLM-assisted novices than for controls on biosecurity knowledge tasks.
  • Evaluation limitations: Evaluation results are sensitive to participant composition, duration, task selection, and other design choices, limiting clear judgments about generalization.

2.4 How Developers Currently Characterize Threat Actors

The surveyed frontier AI safety policies vary substantially in whether and how concretely they characterize threat actors. This heterogeneity matters because evaluation design depends on actor-specific assumptions.

  • Current practice: The surveyed policies differ both in which actor dimensions they highlight and in how concretely they describe actors.
  • Current practice: Policies range from detailed profiles to labels such as “low skilled” or “well-resourced,” categorical identities, or no actor specification.Some policies specify education, time horizon, and financial capacity, while others leave the actor implicit.
  • Evaluation implications: A descriptive persona provides a referent for calibrating evaluations, whereas a categorical label or implicit assumption does not.
  • Evaluation implications: Realism is actor-relative, so underspecified threat actors turn realism from a checkable criterion into a free-floating aspiration.
  • Evaluation implications: Important unresolved questions include what counts as expertise or adequate resources and how evaluation trials reflect actors’ time and financial capacities.
  • Open-weight models: Open-weight developers have had limited public threat-actor disclosure despite facing substantial post-release exposure.

Related Work

Existing work offers useful ways to analyze adversaries, but behavioral frameworks are largely retrospective and actor-profiling efforts remain fragmented and domain-specific. The paper positions standardized prospective characterization as a missing capability for frontier AI misuse evaluations.

  • Risk identification: Risk identification includes understanding the sources of risks and the scenarios through which they are realized, making adversary reasoning relevant to AI evaluations.
  • Behavioral frameworks: Cybersecurity frameworks such as the Diamond Model, ATT&CK, ATLAS, and Kill Chain organize observable behavior or incidents rather than prospective classes of actors.
  • Behavioral frameworks: Retrospective behavioral frameworks are less suited to pre-release reasoning about novel harms and populations that have not yet acted.
  • Actor profiling: Other profiling efforts examine dimensions such as group structure, AI and biological capabilities, operational capacity, or financial resources, but remain scattered or domain-specific.
  • Actor profiling: STIX provides threat-actor fields for type, role, and sophistication, but these are primarily metadata labels rather than an evaluation-calibration scale.
  • Standardization gap: The paper identifies a gap in interoperable standardization across CBRN and offensive cyber AI misuse risk and proposes six attributes as a common language.
  • Standardization gap: Explicit actor calibration is presented as necessary for interpreting evaluation results and comparing, aggregating, and acting on risk analyses.

4.1 Taxonomy Design Goals

The taxonomy is designed to be comprehensive, generalizable, systematic, and operationalizable across CBRN and offensive cyber AI misuse risks. Its matrix structure makes assumptions explicit while allowing attributes to vary independently.

  • Design goals: The framework is intended to be generalizable, fine-grained, comprehensive, systematic, and operationalizable.
  • Generalizability: Its attributes are derived from concerns recurring across CBRN and cyber threat analysis, supporting a common language across frontier AI contexts.
  • Systematic specification: A matrix with explicit fields makes omissions visible and functions as a structural analogue to pre-analysis plans.
  • Operationalizability: Tiered observable criteria translate actor profiles into checkable evaluation assumptions, including realistic elicitation methods and task time constraints.
  • Modularity: Independent attribute variation can expose edge cases that holistic characterizations might miss.

4.2 Attribute Derivation

The paper derives a six-attribute taxonomy from recurring themes in terrorism, biosecurity, cybersecurity, and related operational-capacity literature. It represents actors as combinations of independently tiered capacities rather than single overall risk classes.

  • Attribute derivation: The taxonomy distinguishes technical sophistication, prior domain knowledge, organizational capacity, operational infrastructure, financial capacity, and time horizon.
  • Attribute derivation: The attributes are distilled through a systematic review using Nickerson et al.’s meta-characteristic, ending conditions, and iterative conceptual or empirical derivation.
  • Matrix representation: A complete actor profile is a vector of tiers across all six attributes, and the combination—not any single tier—matters for characterization.
  • Matrix representation: The matrix makes attribute combinations explicit and comparable, with conceptual derivation used for most attributes and empirical derivation for financial capacity and time horizon.
  • Technical sophistication: Technical sophistication helps distinguish misuse risks accessible through existing methods from risks contingent on systematic probing, adaptation, and capability chaining.
  • Technical sophistication: Technical sophistication tiers progress from existing-tool use and prompt interaction through custom workflows, direct weight manipulation, and frontier capability development.
  • Technical sophistication: The resulting tiers are roughly cumulative heuristics, ranging from directing models through interfaces to building pipelines, fine-tuning weights, and developing novel capabilities.

5.2 Prior Domain Knowledge

Prior domain knowledge is modeled as a distinct threat-actor attribute because it determines how much capability an AI system must supply and what assistance is most relevant. Its tiers combine developmental skill models with explicit and tacit knowledge dimensions, and are intended to apply across CBRN and cyber harms.

  • Reasoning Behind Tiers: The taxonomy’s domain-agnostic knowledge tiers apply across CBRN and cyber harm categories while distinguishing explicit and tacit knowledge.
  • Why It Matters: Prior domain knowledge determines how much of an actor’s capability gap AI must bridge and how decisively a release changes what they can achieve.Actors with deep expertise may need marginal assistance, whereas actors lacking expertise may depend on AI for tasks experts consider routine.
  • Why It Matters: An actor’s competency level helps identify whether AI would function primarily as a collaborator, planning aid, or knowledge-gap filler.
  • Reasoning Behind Tiers: The prior-knowledge tiers are grounded in Dreyfus’s skill-acquisition model and Collins’s taxonomy of weak, somatic, and communal tacit knowledge.The framework treats tacit knowledge as qualitatively distinct from explicit, transmittable content.
  • Reasoning Behind Tiers: Ouagrham-Gormley distinguishes novices, sub-experts, and experts by combining theoretical knowledge, disciplinary expertise, and practical weapons-development experience.The WDMP framework adds nested levels from general domain knowledge through expert disciplinary and weapons-specific knowledge.

5.4 Operational Infrastructure

Operational infrastructure captures the physical and digital assets an actor already controls, organizing them by successive breaks in accessibility, control, and scale. These constraints define which attack-chain steps are immediately actionable and shape AI’s marginal contribution.

  • Reasoning Behind Tiers: The infrastructure taxonomy models layered capability ceilings through accessibility, control, and scale rather than cataloging every possible resource.
  • Operational Infrastructure Tiers: General actors use publicly available resources, while Privileged actors can access specialized or restricted resources through institutional or other means.Examples include consumer devices and open-source software for General actors, versus laboratories, regulated materials, and institutional compute for Privileged actors.
  • Operational Infrastructure Tiers: Dedicated actors own and operate purpose-built infrastructure, distinguishing full control from privileged access mediated by a third party.Examples include private laboratories, custom toolchains, dedicated compute clusters, air-gapped systems, and procurement channels.
  • Operational Infrastructure Tiers: Extensive actors operate large-scale, sustained infrastructure that supports complex, parallel operations, including containment facilities, cyber operation centers, and frontier-scale compute.
  • Why It Matters: Infrastructure captures existing control, whereas financial capacity captures what an actor can acquire, so the two attributes constrain operational reach differently.Existing assets determine which attack-chain steps are immediately actionable versus which require additional acquisition.

5.6 Time Horizon

Time horizon measures how long an actor pursues a harmful objective, including preparation, and is calibrated using physical-attack planning data and cyber-intrusion dwell-time data. Longer horizons expand opportunities for refinement, reconnaissance, recovery, and learning, while AI may compress timelines.

  • Why It Matters: Time horizon captures the duration of harmful activity, including preparatory work and the persistent engagement needed for complex multi-step tasks.
  • Reasoning Behind Tiers: Physical-attack data show that roughly 5% of cases involved less than one day of planning, 20% planned within ten days, and 98% concluded within three years.The interquartile range ran from approximately 20 to 95 days.
  • Reasoning Behind Tiers: Cyber-intrusion data report a 14-day global median dwell time in 2025, with 42% of intrusions lasting within one week.The remaining reported bands were 20% for 8–30 days, 27% for 31 days to six months, 6% for six months to one year, and 6% beyond one year.
  • Limitations: The taxonomy uses Smith et al.’s preparation data and Mandiant’s dwell-time data as relative proxies, not precise measurements of equivalent phenomena.The sources measure different phases: pre-incident preparation versus active intrusion after initial access.
  • Reasoning Behind Tiers: Because AI may compress attack timelines, the tiers include granular sub-day and one-to-seven-day categories for AI-assisted operations.

5.7 Assessing Taxonomy Quality

The taxonomy is evaluated as research infrastructure that converts broad adversary descriptions into explicit, populated constraints for evaluation design. Two contrasting profiles show that actors with similar overall capability can require different evaluation priorities.

  • Assessing Taxonomy Quality: Its quality assessment follows Nickerson et al.’s objective and subjective ending conditions, including mutually exclusive, collectively exhaustive, concise, robust, comprehensive, extendible, and explanatory characteristics.
  • Taxonomy Development: The taxonomy was developed by deriving six attributes from CBRN and offensive-cyber literature, distinguishing tiers conceptually or empirically, and reviewing the resulting gradients for stability.
  • Application: A populated threat-actor profile makes the specific combination of constraints and capabilities visible for calibrating evaluation design.
  • Example Profiles: The CBRN profile combines strong domain knowledge and time horizon with low technical sophistication and financial capacity, making collaborator or planning-aid evaluations more pertinent than knowledge retrieval.
  • Example Profiles: The cyber profile combines moderate technical and financial capacity with constrained domain knowledge and a short horizon, favoring knowledge-gap, reconnaissance, and multi-session evaluations.
  • Application: Despite broadly similar overall capability, the two profiles warrant different evaluation priorities; describing both as moderately capable would obscure those differences.

6.2 Evaluation Design Considerations

The taxonomy translates explicit threat-actor assumptions into evaluation design choices across elicitation, domain knowledge, organizational capacity, infrastructure, budget, and time horizon. This mapping is intended to improve calibration and cross-study comparability while avoiding systematic underestimation of adversaries.

  • Technical Sophistication: Evaluations should match elicitation effort to the modeled technical-sophistication tier, from prompt-based jailbreaks for Basic tiers to more advanced methods for Expert tiers.
  • Prior Domain Knowledge: Domain-specific contextual nudges can approximate an adversary’s prior knowledge and avoid confounding model assistance with steps the actor would already know.
  • Organizational Capacity: Single-session, single-actor evaluations may underrepresent organizationally enabled adversaries, motivating sustained multi-session testing with division of labor.
  • Operational Infrastructure: Evaluators should explicitly map available software, scaffolding, and physical equipment to taxonomy tiers so environmental assumptions are declared and comparisons remain faithful.
  • Financial Capacity: Relatively trivial evaluation budgets may systematically underestimate actors in moderate-to-high financial-capacity tiers, so misuse evaluations warrant serious scoping and budgeting.
  • Time Horizon: Insufficient time relative to the modeled time-horizon tier can underestimate adversary capability because participants cannot reach full operational effectiveness.

6.3 Toward Institutional Practice

The taxonomy is proposed as institutional infrastructure for documenting, verifying, and comparing threat-actor assumptions across frontier AI evaluations. The paper envisions its use in safety policies and model cards, with auditable registries that preserve pre-evaluation profiles while accommodating disclosure limits.

  • Institutional Foundation: The taxonomy and its examples provide a conceptual foundation and common language for describing potential threat actors and downstream evaluation use.
  • Standardization Analogy: The taxonomy is framed as infrastructure analogous to ISO 9000’s standardizing role for quality-management systems, while defining six adversary-relevant attributes and 29 terms.
  • Frontier Safety Policies: Frontier AI companies should use all six attributes in frontier safety policies to detail the actors against which evaluations are calibrated.
  • Model/System Cards: Model or system cards should explain how each evaluation attended to the six attributes, linking reported study design to its threat-actor profile.
  • Verification: Independent bodies could audit whether taxonomy attributes are addressed in safety policies and model cards and whether companies conform to their documented setup.
  • Registry: A publicly viewable registry could timestamp profiles before results are known and organize studies by attribute and tier to show coverage gaps.
  • Scope Boundaries: The proposed registry should redact some dual-use information, and the taxonomy does not include separate pre-assessment or periodic recertification steps.
  • Open-Weight Risk: Open-weight releases make prospective adversary reasoning especially urgent because weights cannot be recalled, safeguards cannot be patched, and usage cannot be monitored after release.
Loading 2608.25361v1…