Source-linked AI summary
ICS Cybersecurity Datasets: A Systematic Meta-Review of Coverage, Evaluation Practice, and Structural Gaps
Konstantinos E. Kampourakis, Vyron Kampourakis, Georgios Kambourakis, Sokratis Katsikas, Stefanos Gritzalis, Mario Rodríguez-Béjar, José Luis Hernández-Ramos
TL;DR
Public ICS cybersecurity datasets underpin evaluation, but prior work had not systematically tested whether the collective corpus supports current evaluation claims. This paper meta-reviews 18 studies, harmonizes 83 datasets through a five-dimensional taxonomy, and audits evaluation practice. It finds structural corpus imbalances alongside limited streaming evaluation, train/test discipline, and reproducibility, motivating coordinated improvements in corpus construction, benchmarking, labeling, and data governance.
Problem
Prior work had not systematically assessed whether the collective public ICS dataset corpus supports current evaluation claims.
Method
The paper conducts a meta-review of 18 studies, harmonizes 83 identified datasets with a unified five-dimensional taxonomy, and audits evaluation practices.
Results
The corpus is structurally skewed toward OT Disruption, isolated Stage 2 activity, and simulated or non-operational evidence, with Level 0 field-device evidence effectively absent.
Takeaways & Limitations
The findings motivate cross-stage corpus construction, temporally structured benchmarking, event-level label standards, and governance frameworks for operational data sharing.
Takeaways & Limitations
The reported distributions describe the corpus identified through included surveys and are not exhaustive estimates of the complete ICS cybersecurity dataset space.
Abstract
from arXiv · showhide
Intrusion detection research in Industrial Control Systems (ICS) heavily depends on public datasets, yet no prior work has systematically assessed whether the collective dataset corpus supports current evaluation claims. This paper addresses this gap through a meta-review of 18 studies between 2019 and 2026, from which 83 ICS, or ICS directly related, cybersecurity datasets are identified, harmonised, and characterised using a unified five-dimensional taxonomy. The taxonomy reveals that the corpus is structurally skewed: 85.5% of datasets concentrate on late-stage OT Disruption tactics, cross-stage IT/OT progression sequences are present in only 8.4% of cases, field-device evidence at Level 0 of the Purdue hierarchy is effectively absent, and operationally sourced data accounts for only 15.7% of the collection. A parallel audit of evaluation practices shows that zero report streaming evaluation, fewer than half apply disciplined train/test partitioning, and only two satisfy reproducibility requirements. Furthermore, a taxonomy-evaluation coupling analysis shows that dataset imbalances constrain the scope and feasibility of several evaluation practices. Based on these findings, we identify three structural imbalances: i) architectural shallowness, ii) progression compression, and iii) cross-domain substitution, and derive a coordinated research agenda which covers cross-stage corpus construction, temporally structured benchmarking, event-level label standards, and governance frameworks for operational data sharing.
1 Introduction
ICS cybersecurity research relies on public datasets, but the dataset ecosystem and evaluation practices remain fragmented and uneven. This meta-review unifies prior evidence, characterizes datasets and evaluations, and develops a forward-looking research agenda.
- Public datasets provide a shared experimental substrate for anomaly detection, attack classification, benchmark comparison, and reproducible evaluation.
- The ecosystem is uneven across sectors, modalities, attack representation, labeling conventions, documentation quality, and evaluation procedures.
- Limited cross-dataset generalization evidence raises doubts about model and technique transferability, while also hindering reproducibility and comparability.
- The meta-review synthesizes recent survey literature across sectors, modalities, and review perspectives into a unified account of the ICS dataset landscape.
- Its analysis combines a standard-anchored five-dimensional taxonomy, corpus-level dataset characterization, evaluation and reproducibility auditing, and a future research agenda.
- The proposed agenda translates identified gaps into recommendations for dataset creators, benchmark curators, and related stakeholders.
2 Methodology
The paper conducts a PRISMA-based meta-review of survey and review evidence on datasets used in ICS cybersecurity research. It searches multiple databases, applies predefined screening and appraisal procedures, and retains 18 studies for analysis.
- The analysis treats included studies as the primary identification unit and extracted datasets as the secondary unit for corpus characterization.
- The review identifies, analyzes, and synthesizes studies addressing dataset types, supported security tasks, represented industrial domains, and field gaps.
- The research questions cover recurring dataset resources, corpus representation across five technical dimensions, prior limitations, and recommended future directions.
- The search combines industrial-context, dataset, cybersecurity, and review-oriented terms across IEEE Xplore, Scopus, Web of Science, and ACM Digital Library.
- 361 records were initially retrieved before systematic duplicate removal, title screening, abstract assessment, and full-text eligibility checks.
- The final corpus comprises 18 studies: 16 meeting survey-oriented eligibility criteria and two additional works identified through citation tracing.
- Quantitative distributions are interpreted as descriptive properties of the identified corpus rather than exhaustive estimates of all ICS cybersecurity datasets.
3 Evidence Base and Prior Surveys
The evidence base comprises 18 recent studies with increasing attention to public ICS datasets, but their thematic scope, analytical breadth, and methodological detail vary substantially. Prior surveys diagnose recurring dataset and evaluation problems more consistently than they establish standardized comparison frameworks.
- Corpus overview: The included studies are concentrated from 2019 onward, showing that systematic reflection on public ICS datasets is relatively recent.
- Survey perspectives: Dataset-centric surveys compare features such as attack diversity, protocol support, class balance, modality, labeling, and realism, whereas other studies treat datasets as IDS artifacts or testbed outputs.
- Analytical breadth: Only a subset of studies provides explicit taxonomies or systematic side-by-side dataset comparisons; others provide descriptive summaries or selective exemplars.
- Research gaps: All studies identify recurring limitations involving realism, class imbalance, attack diversity, or evaluation consistency.
- Synthesis: The evidence base supports identification of recurring datasets and shared concerns, but remains inconsistent in terminology, analytical rigor, and thematic focus.
- Survey perspectives: The literature combines dataset-centric, IDS- and method-centric, and testbed- or domain-focused perspectives that examine different aspects of the dataset ecosystem.
- Dataset deficiencies: Prior comparisons report persistent reliance on simulated or small-scale testbeds, narrow sector and protocol coverage, and limited reproducibility.
4 Taxonomies in Related Work
Prior ICS/OT taxonomies span dataset, IDS, attack, and testbed perspectives, often overlapping across classification dimensions. Dataset-centric and IDS-centric perspectives dominate, while attack modeling and testbed design receive less primary attention despite shaping realism and evaluation validity.
- Dataset-centric: Dataset-centric taxonomies classify benchmarks by source, collection method, protocols, attack representation, and realism.
- IDS-centric: IDS-centric taxonomies organize detection approaches by strategy, learning paradigm, model architecture, and operational characteristics.
- Attack-centric: Attack-centric taxonomies classify adversarial behavior by objectives, techniques, or lifecycle stages to assess attack coverage and underrepresented threats.
- Testbed-centric: Testbed-centric taxonomies describe cyber-physical environments through realism, virtualization, process representation, protocol support, and reproducibility.
- Cross-family overlap: The reviewed studies frequently span multiple taxonomy families, linking datasets, attacks, IDSs, and testbeds as connected ecosystem components.
- Coverage imbalance: Dataset-centric and IDS-centric perspectives dominate the literature, whereas fewer studies focus primarily on attack modeling or testbed design.
5 Unified Taxonomy
Prior taxonomies address attack representation, system scope, provenance, and evaluation quality inconsistently, leaving no comprehensive, operationally grounded scheme for comparing the full dataset corpus. The proposed five-dimensional taxonomy anchors classification in recognized ICS frameworks and separates adversary coverage, system scope, provenance, and evidence quality.
- Existing taxonomies operationalize attack representation, system scope, provenance, and evaluation quality differently and rarely address them simultaneously.
- No existing taxonomy is simultaneously comprehensive, operationally grounded, and consistently applied across the full dataset corpus.
- The unified taxonomy grounds each classification dimension in recognized reference frameworks with defined domain vocabulary, boundaries, and groupings.
- D1 and D2 measure adversary tactic coverage and attack progression, including whether datasets capture Stage 1, Stage 2, or continuous cross-stage activity.
- D3 characterizes architectural evidence using Purdue levels, while D4 captures OT environment and dataset provenance.
- D5 addresses operational consequences through evidence quality, while the combined taxonomy supports standard-anchored, like-for-like dataset comparison.
6 Comparative Analysis of Datasets
Across 83 dataset rows, the corpus is structurally uneven: it emphasizes late-stage disruption, Stage 2 scenarios, supervisory-layer evidence, laboratory-derived provenance, and fragmented labels. The resulting taxonomy exposes architectural shallowness, progression compression, and cross-domain substitution that constrain dataset suitability and comparison.
- 83 dataset rows form a structurally uneven corpus dominated by disruption-oriented, Stage 2-centric, and laboratory-derived evidence.
- Adversarial Coverage and Progression: 85.5% of datasets include OT Disruption, while Reconnaissance and Staging appears in 46 datasets and Impact in 20.
- Adversarial Coverage and Progression: 57.8% of datasets represent Stage 2 scenarios, whereas only 8.4% capture cross-stage progression across the IT/OT boundary.
- Architectural Evidence and Purdue-Layer Coverage: Among 41 applicable datasets, 18 are coded at Levels 1–2 and 19 at Levels 2–3; only four reach Levels 3–4, with no explicit Level 0 evidence.
- Operational Realism and Data Provenance: 39.8% of datasets are physical-testbed provenance, compared with 15.7% operational, 16.9% emulated or virtualized, and 12.0% simulation-only.
- Detection Evidence Quality: D5 labels are fragmented across meaning, granularity, comparability, and evidence units, including 25 flow-level, 24 record-level, and 13 packet-level datasets.
- Structural Imbalances: IoT and general IDS datasets can support general intrusion-detection research but cannot close gaps in OT-specific architectural evidence.
7 Evaluation Practices
Evaluation practices across the 18 surveyed studies are heterogeneous and frequently offline, pointwise, single-dataset, and vulnerable to temporal leakage. The audit further shows that evaluation validity is coupled to dataset coverage, architecture, provenance, and ground-truth quality.
- Evaluation reporting is heterogeneous across partitioning, metrics, threshold calibration, and related methodological choices.
- Data Partitioning and Temporal Leakage: Random record-level splits and standard k-fold cross-validation can place temporally adjacent samples from one process scenario in both training and testing.
- Data Partitioning and Temporal Leakage: Temporal leakage can inflate F1-score and AUC relative to chronological partitioning, with larger gaps under stronger autocorrelation and larger feature windows.
- Ground-Truth Representation and Label Granularity: Event-aware evaluation remains rare because studies often reduce attack timestamps to record-level accuracy and F1, overlooking detection timing.
- Metrics and Operational Alignment: Alarm rate per unit time and alert burden per event are rarely and inconsistently reported across the surveyed corpus.
- Generalization Evidence and Transfer Validity: None of the surveyed works report streaming or pseudo-online evaluation approximating deployed ICS operation.
- Taxonomy–Evaluation Coupling: D3, D4, and D5 constrain transfer, deployment-oriented, and event-aware evaluation through limited architectural evidence, provenance realism, and label precision.
- Taxonomy–Evaluation Coupling: Credible progress requires coordinated improvement of datasets and evaluation methodology rather than isolated changes at either level.
8 The Road Ahead
The review identifies structural imbalances in the ICS dataset corpus and shows that dataset limitations constrain evaluation practice. It proposes coordinated advances in dataset construction, evaluation protocols, labeling, benchmarking, governance, and shared infrastructure.
- Structural diagnosis: 83 datasets reveal three interlocking deficits: architectural shallowness, progression compression, and cross-domain substitution.These imbalances constrain what the research community can claim about ICS intrusion detection.
- Architectural and provenance coverage: 15.7% of datasets are operationally sourced, compared with 39.8% physical-testbed coverage and 28.9% combined simulation/emulation.Federated sharing and hybrid provenance designs are proposed to increase operational relevance while addressing disclosure constraints.
- Labeling and comparability: 89.2% of datasets use live-executed labels, but event-level onset and offset annotations remain uncommon and multi-class schemes are rarely aligned.A community label schema with temporal boundaries and sensor-level attribution is proposed to improve cross-dataset comparability.
- Evaluation and infrastructure: None of 18 studies report streaming evaluation, only 7 apply adequate partitioning discipline, and only 2 fully report reproducibility.The coupling analysis links scarce cross-stage data to cross-stage evaluation limits, limited operational data to streaming constraints, and coarse labels to event-aware metric limitations.
- Research agenda: Transfer-oriented benchmarks, temporally structured protocols, shared registries, benchmark suites, and governance frameworks are proposed as coordinated priorities.Recommended protocols include timestamp-preserving formats, non-stationary splits, detection delay, alarm burden, and shared preprocessing and dataset splits.
- Research agenda: Cross-stage dataset construction should combine digital twins, cyber ranges, co-simulation, and instrumented red-team exercises to capture end-to-end IT/OT sequences.The proposed approaches target the architectural boundary separating Stage 1 activity from Stage 2 physical effects.
9 Conclusion
The paper asks whether public ICS cybersecurity datasets support current evaluation claims and answers negatively in several measurable respects. Through a systematic characterisation and evaluation audit, it finds structural dataset skew and links those gaps to constrained evaluation validity, motivating coordinated improvements.
- The paper examines whether the collective body of public experimental data supports evaluation claims in ICS cybersecurity research.
- 83 datasets identified across 18 studies are characterised using a standard-anchored D1–D5 taxonomy, with distributions, an evaluation audit, and a coupling map.
- 85.5% of datasets concentrate on OT Disruption, 57.8% capture Stage 2 in isolation, only 8.4% span multiple stages, Level 0 evidence is effectively absent, and operational data accounts for 15.7%.
- The coupling analysis links dataset scarcity and coarseness to limits on cross-stage, streaming, concept-drift, delay-aware, and alarm-burden evaluation.
- Dataset adequacy and evaluation validity must be assessed jointly because architecturally shallow or Stage 2-only data cannot support corresponding detection claims.
- The paper calls for coordinated progress in dataset construction, evaluation methodology, shared benchmarking, and governance for operational data sharing.
A Search Queries Used in Each Database
The meta-review searches were executed on June 18, 2026, using database-adapted queries with a shared conceptual structure. Each query combined four blocks covering ICS/OT domains, datasets and benchmarks, cybersecurity, and review-oriented publications.
- June 18, 2026 was the last execution date for the database searches.
- The queries were adapted to each database's syntax while preserving the same conceptual structure.
- Four query blocks covered industrial-control and operational-technology environments, datasets and benchmarks, cybersecurity, and review-oriented publications.
A.1 Web of Science
The Web of Science query required records to match ICS or related infrastructure terms, dataset or benchmark terms, cybersecurity terms, and review-oriented publication terms.
- The Web of Science query combined industrial-control, dataset, cybersecurity, and review-related term groups.
- Its domain block included industrial control system, ICS, SCADA, OT, IIoT, operational technology, and critical infrastructure.
- Its publication block included survey, review, meta-analysis, systematic review, and literature review.
A.2 IEEE Xplore
The IEEE Xplore query applied the same four conceptual blocks through abstract-field searches. It required abstract matches for domain, dataset, cybersecurity, and review terms.
- The IEEE Xplore query searched abstract fields for ICS, SCADA, OT, IIoT, operational technology, and critical infrastructure terms.
- It required abstract matches for dataset, benchmark, testbed, or data collection terms and for cybersecurity, intrusion detection, anomaly detection, attack, or IDS terms.
- Its review block required abstract terms including survey, review, meta-analysis, systematic review, or literature review.
A.3 ACM Digital Library
The ACM Digital Library search used abstract-field syntax for the same domain, dataset, cybersecurity, and review concepts, alongside a TITLE-ABS-KEY formulation. The appendix also identifies the manuscript as submitted to ACM.
- Abstract-field query: The first ACM query used bracketed Abstract fields for industrial-control, dataset, cybersecurity, and review-related term groups.
- Abstract-field query: Its domain terms included industrial control system, ICS, SCADA, OT, IIoT, operational technology, and critical infrastructure.
- Abstract-field query: Its dataset, cybersecurity, and review blocks included benchmarks and testbeds, intrusion or anomaly detection and attacks, and survey or systematic-review terms.
- TITLE-ABS-KEY query: A second ACM formulation used TITLE-ABS-KEY for the same domain, dataset, cybersecurity, and review concepts.
- The appendix is associated with a manuscript submitted to ACM.