Source-linked AI summary

A Case-Control Measurement Study of OSINT Source Effectiveness for Critical Infrastructure Defense

Ekrem E. Emeksiz, Jeel Piyushkumar Khatiwala, Divyangkumar Patel, Weifeng Xu

arXiv:2608.21471v1cs.CR

TL;DR

Critical-infrastructure defenders lack empirical evidence about which public OSINT sources precede attacks and therefore must prioritize among many feeds. This paper analyzes 54 confirmed attacks and 12 null cases to classify ten source classes by coverage, contamination, and lead time, finding three operational profiles and high coverage from small portfolios.

  • Problem

    Critical-infrastructure defenders lack empirical evidence for evaluating public OSINT source effectiveness and selecting limited source portfolios that maximize confirmed-attack coverage.

  • Method

    The study audits ten public OSINT source classes across 54 confirmed attacks and 12 null cases, measuring attack coverage, null-case contamination, signal lead time, and portfolio performance.

  • Results

    Sources separate into three operational profiles: six precursor classes with zero null firings, three disclosure-exposure classes, and one broad-coverage class with 91.3% within-corpus precision.

  • Takeaways & Limitations

    Two sources cover 92.6% of corpus attacks and three cover 96.3%, while disclosure-exposure sources belong in a patch-management queue rather than incident response.

  • Takeaways & Limitations

    The corpus has limited per-source null-side power and includes only attacks with at least one observable pre-attack signal, so absolute coverage figures may decline for silent attacks.

Abstract

from arXiv · show

Defenders of critical infrastructure (CI) subscribe to many public open-source intelligence (OSINT) feeds without an empirical basis for which feeds actually precede attacks. We provide one. Across 54 confirmed CI cyberattacks from 2010 through 2024 spanning twelve named CI sectors plus a cross-sector category (consolidation rules in Section IV), paired with 12 null-control vulnerability cases drawn from the same source space, we audit per-source attack coverage, null-case contamination, and signal lead time for ten public OSINT source classes that meet a minimum-volume threshold. Sources separate cleanly into three operationally distinct mission profiles (pooled Fisher exact p = 3.4x10^-8): precursor (six classes with zero observed null firings at coverage at or above 5%), disclosure-exposure (three classes whose null contamination meets or exceeds attack coverage), and one large broad-coverage class that mixes the two profiles but retains 91.3% within-corpus precision. The precision-side classification is stable across a 2019 temporal partition and across a US-versus-non-US geographic partition. Two sources, one broad-coverage and one precursor, cover 92.6% of corpus attacks; three cover 96.3%. The greedy portfolio at k = 3 outperforms the mean random three-source subset by 39.8 percentage points. Several source classes widely treated as canonical for industrial control system defense fall into the disclosure-exposure profile by operational mission, not by quality. Per-sector, per-actor, and per-jurisdiction portfolios diverge in rank order despite a shared rank-one source. The corpus, linkage protocol, and classification rules are released.

I. INTRODUCTION

The paper addresses the lack of empirical evidence for which public OSINT sources precede confirmed critical-infrastructure attacks. It uses paired attack and null-control cases to measure source behavior and select constrained source portfolios.

  • I. INTRODUCTION: 54 confirmed attacks and 12 null-control vulnerability cases form a fifteen-calendar-year corpus for auditing OSINT source firings, lead time, and contamination.The methodology adapts case-control study design from vulnerability exploitation analysis to OSINT source evaluation.
  • I. INTRODUCTION: Case-control pairing enables source precision measurement rather than attack coverage alone.The null cases are drawn from the same source space as the attacks.
  • I. INTRODUCTION: Three operational source profiles separate by joint attack coverage and null contamination with pooled Fisher exact p = 3.4 × 10−8.The profiles are precursor, disclosure-exposure, and broad-coverage.
  • I. INTRODUCTION: 77.8% coverage, 33.3% null contamination, and 91.3% precision characterize the vendor-tier research broad-coverage class.The precision-side classification remains stable across temporal and geographic partitions.
  • I. INTRODUCTION: 92.6% of corpus attacks are covered by a two-source portfolio, while three sources cover 96.3%.The portfolio combines one broad-coverage source with one precursor source at two sources.
  • I. INTRODUCTION: The released OSINTCI-66 corpus contains 15 years, 54 attacks, 12 nulls, and 161 verified signals across 15 observed source classes.Construction used a two-verifier-with-adjudication protocol and provides a full audit trail.

II. BACKGROUND AND RELATED WORK

Prior CTI research measures feed behavior and content, but does not address source-level attack precedence against null controls. This paper defines source metrics and a cardinality-constrained portfolio objective for defenders operating under finite attention.

  • II. BACKGROUND AND RELATED WORK: Existing CTI studies measure precision-recall variance, feed overlap, TTPs, and quality, but do not test source-level attack precedence against null controls.Related prediction work generally operates per-CVE, whereas this study operates per-source.
  • II. BACKGROUND AND RELATED WORK: A pre-attack signal is a publicly observable data point published before a confirmed attack with verifiable linkage via CVE, malware family, threat actor, or infrastructure.Coverage counts any verified-linkage signal, while lead time is reported separately.
  • II. BACKGROUND AND RELATED WORK: The defender selects a source portfolio under finite monitoring attention to maximize the fraction of future attacks covered by at least one selected source.The threat model includes SOCs, threat-intelligence teams, and CISO offices facing ransomware, nation-state, and hacktivist actors.
  • II. BACKGROUND AND RELATED WORK: Source precision is the within-corpus fraction of incidents fired on by a source that are confirmed attacks.Null contamination is the analogous coverage measure over the null corpus.
  • II. BACKGROUND AND RELATED WORK: Source selection is formulated as cardinality-constrained maximum coverage, with binary variables for monitored sources and reached attacks.The coverage function is monotone submodular, so greedy optimization has a (1−1/e) ≈63% guarantee under the cardinality constraint.
  • II. BACKGROUND AND RELATED WORK: Within-corpus precision supports source ranking but is not deployment-time positive predictive value because deployment base rates differ from the corpus 54:12 ratio.This distinction limits direct interpretation of precision for operational deployment.

IV. EMPIRICAL AUDIT

The audit constructs and processes OSINTCI-66, linking public signals to 54 confirmed attacks and 12 null-control cases to measure source coverage, contamination, and lead time.

  • Pipeline: Nine processing stages run from case selection through signal collection, linkage, verification, cleaning, aggregation, metric computation, statistical analysis, and portfolio optimization.The pipeline is specified across Sections IV–VII, with metrics defined in Section III.
  • Corpus: 54 confirmed CI cyberattacks and 12 null-control cases form the OSINTCI-66 corpus across twelve named sectors plus a cross-sector category.The corpus spans May 2010 through October 2024 and applies stated sector-consolidation rules before stratification.
  • Linkage: Signals qualify as linked when they satisfy at least one of seven criteria, including CVE, malware-family, threat-actor, or infrastructure matches.The strict four criteria are used for headline results.
  • Verification: Two independent verifiers classify outcome rows as clean-pass or clean-fail, while disagreements undergo adjudication or team discussion.Across 251 outcome rows, pre-adjudication agreement was 84.5%.
  • Metrics: Precision is computed from attack and null firing counts, with 54 attacks and 12 null cases as the corpus denominators.The 95% lower bound uses a joint Clopper-Pearson construction, while median lead uses the proper sample median over nonnegative leads.

B. Per-source measurement

Per-source measurement reports verified signal counts, attack and null coverage, precision with uncertainty bounds, median lead time, and assigned mission profile.

  • Per-source metrics: Table I reports verified signal counts, covered attack incidents, covered null cases, source precision, median attack-signal lead time, and mission profile for each class.Precision includes Clopper-Pearson 95% lower bounds.
  • Uncertainty: With zero null firings over 12 nulls, point precision remains 100%, but one-sided 95% bounds on true contamination reach 22.1%.The study therefore treats coverage-dependent lower bounds, ranging from 23.8% to 78.5%, as the honest summary.

V. SOURCE MISSION TAXONOMY

Joint attack-coverage and null-contamination measurements separate sources into precursor, disclosure-exposure, and broad-coverage operational profiles.

  • Precursor: Six precursor classes show zero observed null firings with attack coverage between 5.6% and 27.8%.US-CERT TA/AA leads at 27.8% coverage with a 32-day median lead; MalwareBazaar and VirusTotal illustrate tighter attack-onset coupling.
  • Disclosure-exposure: Three disclosure-exposure classes have null contamination meeting or exceeding attack coverage and primarily serve patch-management functions.NVD has 50.0% precision and 91.7% null contamination; ICS-CERT has 33.3% precision and 66.7% null contamination.
  • Figure 1: Figure 1 places Vendor-Tier1 above the 50% base-rate-balanced precision diagonal and disclosure-exposure classes below it.The diagonal is defined by nct=cov.
  • Broad-coverage: Vendor-Tier1 is the broad-coverage class, covering 77.8% of attacks with 33.3% null contamination and 91.3% within-corpus precision.The class aggregates Tier-1 vendor and major incident-response research; decomposition would separate precursor and disclosure-exposure subsets.
  • Statistical separation: Pooled precursor firings are 39 attacks versus 0 nulls, while disclosure-exposure firings are 24 attacks versus 23 nulls; the one-sided Fisher test gives p = 3.4 × 10^-8.At incident granularity, at least one precursor fires on 31 of 54 attacks and 0 of 12 nulls, with p = 1.7 × 10^-4.
  • Validation: The taxonomy’s within-class validation finds significant deviations for Vendor-Tier1, NVD, and ICS-CERT after Benjamini-Hochberg correction.These are the three classes surviving correction among classes with at least two combined firings.

VI. STABILITY UNDER TEMPORAL AND GEOGRAPHIC PARTITION

The source-mission classification remains stable across temporal and geographic partitions, while lower-ranked sources and portfolio choices vary with the partition.

  • Validation design: The study treats the taxonomy as descriptive rather than predictive and tests it with out-of-sample temporal and geographic partitions.This addresses the concern that deriving and evaluating a taxonomy on the same data can mask heterogeneity.
  • Temporal partition: A 2019 split preserves Vendor-Tier1 as rank-one greedy source and preserves the disclosure-exposure labels for NVD, ICS-CERT, and CISA-KEV.The pre-2019 subset contains 11 attacks versus 43 from 2019 through 2024, and precursor ranks shift with ecosystem timing.
  • Geographic partition: Vendor-Tier1 remains rank one in both US and non-US subsets, covering 83% of US-target attacks and 73% of non-US attacks alone.The second source differs: US portfolios use CISA-KEV or US-CERT, while the non-US portfolio reaches 97% at k=3 with CERT-UA.

A. Greedy frontier and exhaustive verification

The greedy frontier selects sources by marginal new-attack coverage under a deterministic tie-break, and exhaustive enumeration confirms the greedy portfolio is optimal at every evaluated size.

  • Greedy selection: Greedy adds the source maximizing marginal new-attack coverage across eight defender-relevant classes and stops when marginal gain reaches zero.The frontier uses broad-coverage, precursor, and borderline CISA-KEV classes.
  • Greedy selection: Ties are resolved deterministically by mission profile, then alphabetical source-class order.Broad-coverage precedes precursor, which precedes disclosure-exposure.
  • Tie resolution: At k=4, either NCSC-Tier1 or CERT-UA yields identical 98.1% coverage, while CISA-KEV is selected at k=5 when its remaining coverage is unique.At k=3, MalwareBazaar is selected over CISA-KEV under the mission-profile tie-break.
  • Verification: Exhaustive enumeration over the eight defender-relevant classes confirms that greedy matches the optimum at every portfolio size.The benchmark evaluates the greedy frontier against exhaustive alternatives.
  • Frontier: 96.3% attack coverage at k=3 compares with 56.5% mean-random and 22.2% worst-case coverage across 35 three-source subsets.The comparison covers subsets of the seven precursor and broad-coverage classes.

B. Commercial collection scope and marginal value

The study measures publicly accessible vendor and incident-response output rather than paid enterprise feeds, finding substantial marginal value for the aggregated Vendor-Tier1 class while restricting recommendations to sufficiently populated strata.

  • Collection scope: Vendor-Tier1 aggregates public research from major vendors and incident-response firms, while paid enterprise feeds were not procured or tested.The class includes public blogs, advisories, and threat reports from named Tier-1 providers.
  • Marginal value: 29.6 pp is Vendor-Tier1’s rank-one marginal coverage above the purely-free portfolio, which reaches 70.4% without it.Removing Vendor-Tier1 and rerunning greedy produces a seven-source ordering that saturates at 70.4%.
  • Stratified portfolios: Cells with fewer than five attacks are excluded from recommendations but retained for recomputation as additional incidents accumulate.The exclusion covers eight of thirteen sector strata and most sector-by-actor cells.
  • Stratified portfolios: Vendor-Tier1 is rank one in every stratum with at least three attacks except Transport, where US-CERT is rank one.The stratification covers sector and actor-type portfolios.
  • Stratified portfolios: Cybercriminal-heavy strata close through CERT-UA, MalwareBazaar, and CISA-KEV, whereas Nation-State strata close through US-CERT under the tie-break.The closing sources track actor mix and associated public samples or exploited-CVE disclosures.

IX. ADAPTIVE ADVERSARY ANALYSIS

Adversary suppression affects portfolio coverage unevenly, while case examples and lead-time analysis show both specialized early warning and near-zero-day limits.

  • Suppression scenarios: 70.4% attack-precursor coverage remains after losing Vendor-Tier1 entirely.This retained coverage comes from sources the adversary cannot structurally suppress.
  • Suppression scenarios: 38.9% coverage remains under purely government-CSIRT channels.This scenario excludes vendor research, malware repositories, community exchanges, and CISA-KEV disclosure.
  • Walked examples: 0 positive lead days are reported for MOVEit Transfer under the same-day NVD inclusion convention.The case qualifies through a same-day NVD entry, while other listed publications were post-attack.
  • Walked examples: 259 days of warning precede Volt Typhoon in the clearest specialized-precursor example.The warning came from a joint advisory and Vendor-Tier1 publication.
  • Lead-time analysis: 23 days is the median lead time across 117 verified-signal occurrences.27.4% appear within 7 days, 56.4% within 30 days, and 87.2% within 180 days; weekly review therefore discards about half the available warning.

XI. DISCUSSION

The discussion presents mission separation, portfolio specialization, and operational boundaries, while highlighting corpus, visibility, and validation limitations.

  • Findings: p = 3.4 × 10^-8 supports a highly significant three-profile source separation that remains stable across temporal and geographic partitions.The structure is also stable across nearly every stratum.
  • Portfolio implications: 92.6% of corpus attacks are covered by two sources, rising to 96.3% with three.The sources combine one broad-coverage class with precursor classes.
  • Portfolio implications: Per-sector and per-actor portfolios diverge at ranks two and three despite sharing Vendor-Tier1 as rank-one anchor.Examples include CISA-KEV for Healthcare/Cybercriminal, CERT-UA and MalwareBazaar for Manufacturing/Cybercriminal, and US-CERT for Energy/Nation-State.
  • Limitations: Coverage is conditional on attacks having an observable public pre-attack signal.Silent attacks contribute no observable signals, so absolute coverage degrades proportionally under uniform silent-attack scaling.
  • Limitations: The study measures publication before attack, not whether defenders observed or acted on the signal.Within-corpus precision also does not model per-day alert volume.
  • Limitations: External validation against post-2024 data and live SOC deployment remains open.The temporal and geographic partitions are within-corpus stability checks rather than external validation.

XIII. CONCLUSION

The paper delivers a null-controlled empirical framework for evaluating CI OSINT sources, selecting portfolios, and releasing an auditable research corpus.

  • Conclusion: p = 3.4 × 10^-8 separates public OSINT classes into precursor, disclosure-exposure, and broad-coverage profiles.The precision-side classification is stable across temporal and geographic partitions.
  • Conclusion: 92.6% of corpus attacks are covered by two sources, and 96.3% by three.The greedy k=3 portfolio beats the mean random three-source subset by 39.8 percentage points.
  • Conclusion: The taxonomy classifies several canonical ICS-defense sources by operational mission rather than quality.Some heavily promoted sources fall into the disclosure-exposure profile.
  • Scope: All data come from publicly accessible OSINT without privileged, hacked, or non-public sources.The taxonomy primarily benefits under-resourced defenders, while adversary suppression is quantified against non-suppressible sources.
  • Released artifact: The released artifact includes incident records, linked signals, null-selection logs, and release-preparation records.Tables and figures reproduce from the CSVs without project code.
Loading 2608.21471v1…