Source-linked AI summary

Research Methodologies for Cybersecurity in Enterprise Environments: A Narrative Review, Synthesis and Executable Guide

Tran Duc Le

arXiv:2608.24850v1cs.CR

TL;DR

Enterprise cybersecurity research lacks a single method suitable for its diverse questions, outcomes and organisational settings. The paper narratively synthesises a verified corpus into eleven methodology families and executable protocols, finding that methodological variation and validity reasoning must govern evaluation while the review’s selected corpus limits claims of completeness.

  • Problem

    Enterprise cybersecurity studies use methods for different outcomes, but researchers may lack a basis for judging when a familiar method is the wrong instrument.

  • Method

    The paper provides a narrative synthesis of a verified corpus, maps eleven methodology families to questions and claim types, and turns each family into an executable protocol.

  • Results

    The synthesis supports explicit validity reasoning and shows that the review’s 151-work corpus describes its selection rather than the composition of the field.

  • Takeaways & Limitations

    Researchers should treat methodological choice as a claim-and-decision matching problem and make selection, verification and validity conditions inspectable.

  • Takeaways & Limitations

    As a narrative review, the corpus is selected rather than exhaustive and reflects the authors’ judgement and search facets.

Abstract

from arXiv · show

Enterprise cybersecurity research draws on a wider range of methods than any single community routinely teaches. Researchers face a selection problem before they face a technical one: a study may simultaneously need a systematic review, a design-science artifact, a controlled detection experiment, an interview study, or an attack-graph model. This paper addresses that problem in two ways. First, it provides a narrative review and synthesis of methodological practices across a verified corpus of 151 works. We organise these practices into eleven methodology families, detailing for each what questions it answers, the strength of its supporting evidence, and its common failure modes. Second, we convert each family into an executable protocol comprising ordered steps, required instruments, evaluation criteria, common validity threats, and a reporting checklist. Every protocol is also visually mapped to make the sequence, decisions, and threats legible at a glance. We also treat contradictions in the literature as evidence. For example, reported rankings of intrusion-detection algorithms are wildly inconsistent across individually careful studies. We argue this pattern is most parsimoniously explained by variations in evaluation design rather than the algorithms themselves, as these studies differ in design dimensions known to shift results by more than the margins separating the algorithms. Ultimately, the evidence supports methodological pluralism disciplined by explicit validity reasoning. We conclude that researchers must match their evaluation design to the decision under study, triangulate technical against organisational evidence, explicitly state the population a result generalises to, and report the conditions under which the result would not hold.

I. INTRODUCTION

Enterprise cybersecurity research spans heterogeneous questions and settings, making methodology selection a central problem rather than a purely technical choice. This paper responds with a narrative synthesis and executable, validity-aware protocols while explicitly limiting claims about coverage and generalisation.

  • Research context: Enterprise cybersecurity research covers technical, organisational and cyber-physical domains where weaknesses can propagate into operational disruption.The scope includes security assessment, penetration testing, threat modeling, attack simulation, incident response, human factors and governance.
  • The selection problem: Studies evaluate different outcomes, so methods trained on one tradition may be poor instruments for another research question.Machine-learning studies emphasise accuracy and F1, while incident-response, risk-modeling and training studies use different outcome measures.
  • Paper response: The paper organises enterprise cybersecurity research into eleven methodology families mapped to research questions and licensed claim types.The taxonomy is paired with protocols, validity threats, evaluation criteria and reporting requirements.
  • Paper response: Each family is converted into an executable protocol with ordered steps, decisions, required outputs, validity threats and a checklist.The protocols and figures are intended to support study design and review in practice.
  • Review design: The review uses a narrative synthesis because its question spans research traditions rather than a commensurable set of primary studies.Its selection and verification procedures are reported explicitly, including source-record checks for identifiers and bibliography metadata.
  • Scope and limits: The corpus contains 151 works from 1977–2026, but its counts describe this review’s selection rather than the field’s composition.The corpus has a median publication year of 2020, with 58 works dating from 2023 onward.

F. Limitations of this review

This narrative review is a defensible, inspectable partition rather than an exhaustive census, with selection, retrieval, appraisal, and coverage limitations. The authors compensate by reporting their procedure, verification record, and known under-coverage.

  • Scope and selection: The corpus is selected rather than exhaustive, and the eleven families should be treated as a defensible partition of what was found.The authors’ security-engineering and empirical-software-engineering background also shaped the space allocated to methodology families.
  • Scope and selection: Facet-by-facet database retrieval means unformulated facets were not covered, while practices that were never published remain invisible.The review also lacked full text for every included work, limiting extraction where only partial material was available.
  • Appraisal: No formal risk-of-bias instrument was applied, so confidence labels are reasoned judgments rather than instrument scores.This limits how directly readers should interpret the review’s confidence assessments.
  • Mitigation: The review reports its retrieval procedure, inclusion criteria, verification record, and known under-coverage to make a non-exhaustive selection inspectable.This is presented as the compensating practice recommended for narrative reviews.
  • Audit: The audit found all 49 references checked out, but rejected an antecedent screening flow whose totals and process narrative were unsupported.A clean reference list and a credible retrieval process require separate checks.
  • Synthesis limits: The synthesis did not pool effects because outcomes such as accuracy, time-to-compromise, maturity level, κ, and perceived usefulness are incommensurable.Comparative-effectiveness conclusions therefore inherit the dataset limitations and reporting inconsistencies of the summarized studies.

E. Failure modes •

The paper identifies recurring failure modes in reviews and design-science studies, especially unverifiable screening, unmeasurable objectives, weak evaluation, and poor reimplementation detail. These failures constrain the claims a method can support.

  • Review failures: A PRISMA flow that does not reconcile, or records no full-text exclusions, signals that screening was not performed as described.The paper treats this as a reader-checkable process failure.
  • Review failures: Vote counting gives studies of different quality and size equal weight, while single-database retrieval inherits one publisher-coverage bias.Both practices can distort a review’s synthesis before interpretation begins.
  • Review failures: Synthesising incommensurable outcomes into an apparent consensus creates a misleading comparison across studies.The paper separately recommends distinguishing review limitations from limitations of included studies.
  • Design-science failures: Design-science objectives written after the build cannot fail, and formative evaluation improves an artifact without testing its effectiveness.A demonstration on one constructed example establishes feasibility only.
  • Design-science failures: Operational-effectiveness claims require rung 4 or 5, but the reviewed corpus evaluates predominantly at rungs 2 and 3.The paper therefore assigns limited confidence to operational-effectiveness claims from design science alone.
  • Reporting failures: Common design-science weaknesses include irreproducible artifact descriptions, collaborator-only expert panels, and no comparison with practitioners’ existing baseline.The reporting checklist requires measurable pre-build objectives, reimplementation detail, named evaluation type, validity threats, and artifact release.

VI. FAMILY C: QUALITATIVE AND SOCIO-TECHNICAL INQUIRY

Qualitative and socio-technical inquiry reveals organisational mechanisms and practitioner perspectives that technical measures miss. Its rigour depends on choosing a reflexive or codebook design before coding and reporting the corresponding standards, sampling, analysis, and limitations.

  • Purpose and use: Qualitative inquiry explains practitioner action, organisational work, and what technical measures omit when variables are unknown or results are inexplicable.It is suited to organisational phenomena and exploratory questions.
  • Socio-technical inquiry: Matched-centre interviews located security-operations problems in relationships between analysts and managers, requiring sampling both sides of the boundary.A complementary study also compared centre performance metrics with what those metrics omit.
  • Analytic procedure: Braun and Clarke’s six phases require active familiarisation, systematic coding across the dataset, and themes understood as shared meaning rather than topic summaries.The framework is widely used but often misapplied, especially through treating themes as emergent or adding unsupported inter-coder agreement.
  • Data collection: Semi-structured interviews use purposive sampling, open questions with planned probes, and a stopping rule whose operationalisation must be stated rather than justified only by “saturation.”Saunders et al. distinguish incompatible meanings of saturation and require researchers to specify which one they used.
  • Rigour and reliability: Reflexive and codebook designs require different rigour criteria: reflexive analysis uses reflexivity, an audit trail, and analytic depth, whereas codebook designs report agreement.The choice must be made before coding because it determines the subsequent rigour standard.
  • Reporting: Qualitative studies should report their design, sampling, recruitment, interview guide, coding, participant attribution, and any agreement statistic with coder count and double-coded proportion.COREQ provides a 32-item reporting checklist for interview and focus-group studies.
  • Evidence and limitations: Confidence is moderate that coordination, cognitive workload, rule conflict, and temporal variation affect security outcomes, but self-report bias, small purposive samples, sector-specific sites, and limited transferability remain limitations.These constraints bound how broadly qualitative findings can be transferred.

VII. FAMILY D: SURVEY, BEHAVIOURAL AND LONGITUDINAL DESIGNS

Survey, behavioural, longitudinal, and Delphi designs answer different population, within-person, and expert-consensus questions. Their validity depends on matching the instrument to the claim, establishing measurement and sampling adequacy, and not treating consensus as evidence of effectiveness.

  • Survey designs: Surveys estimate construct distributions and covariation in defined populations, especially for latent constructs such as perceived threat severity and security culture.Common theoretical frames include protection motivation and institutional governance accounts of policy compliance.
  • Survey protocol: Survey protocols separate population from sampling frame, validate constructs, size samples for the intended analysis, pilot 10–30 target participants, and report participation and completion rates.Measurement reliability, AVE, and HTMT should precede structural paths, R2, and f 2.
  • Generalisability: Convenience-sample results transfer differently by measure and demographics, so generalisation requires stating the target population and testing non-response or coverage concerns.Redmiles et al. compared census-representative, crowdworker, and web-panel results and found that some estimates transferred acceptably while others did not.
  • Longitudinal designs: Cross-sectional surveys capture between-person patterns at one time and cannot identify within-person patterns, making them unsuitable when theory concerns behavioural change within individuals.A four-week experience-sampling study found cybersecurity behaviour varied within the same person over time.
  • Delphi designs: Delphi is appropriate when questions require structured expert judgement rather than a representative employee sample, using successive anonymous rounds that aggregate and rerate prior responses.Campbell’s example used 20 information-security practitioners across three rounds to produce prioritised areas and converged countermeasures.
  • Claim discipline: Delphi licenses a claim about qualified expert consensus, not that the proposed countermeasures work.Reporting should include the expertise criterion, panel size, attrition, stopping rule, and consensus statistic.
  • Evidence and limitations: Survey-design confidence is moderate conditional on measurement-model reporting, while confidence in the within-person finding is limited because it rests on one four-week study.The methodological case for experience sampling is independent of that study’s result.

VIII. FAMILY E: THREAT MODELING, ATTACK GRAPHS AND QUANTITATIVE RISK

Threat modeling and quantitative risk methods connect system structure, attack paths, probabilities, investment, and control prioritisation to enterprise decisions. Their usefulness is moderate but generalisation is constrained by expert assumptions, limited real-incident validation, and changing threat conditions.

  • Threat modeling protocol: Threat modeling decomposes systems into processes, data stores, flows, external entities, and trust boundaries before enumerating applicable STRIDE categories.The data-flow diagram is the working artifact, and element-to-category mapping makes the method systematic rather than free-form.
  • Threat modeling limits: STRIDE supports lightweight component-level analysis but can miss threats arising only from interactions between components.Mitigation analysis should accompany enumeration so the resulting threats can inform action.
  • Attack paths: STRIDE enumerates threats but does not compose multi-step attacks, which require path representations such as attack trees or attack graphs.Attack-graph formalisms can simulate attack sequences and time-to-compromise rather than enumerate them manually.
  • Quantitative risk: Probabilistic architecture models estimate successful-attack probabilities, while decision-theoretic models jointly represent defensive and recovery investment through expected utility.Investment models require stated sources for p(x), L(y), and risk attitude, plus sensitivity analysis because the optimum depends on assumptions.
  • Network analysis: Network analysis uses betweenness centrality to identify techniques recurring structurally across enterprise attack pathways and support control-prioritisation ordering.Centrality reflects taxonomy construction and documented-procedure bias unless edges are weighted by observed incident frequency.
  • Evidence and validity: Threat-modeling and graph-based risk analysis receive moderate confidence, but case-specific assumptions, expert dependence, limited real-incident validation, and changing threats reduce generalisation.Common failures include treating CVSS base scores as risk estimates, validating with the same experts who supplied parameters, and omitting sensitivity analysis.

IX. FAMILY F: SIMULATION AND MODEL-BASED ANALYSIS

Simulation is suited to exploring configuration changes, cascading consequences, and parameter sensitivity when direct experimentation is unsafe or impractical. Its conclusions depend on explicit model specification, parameter provenance, verification, validation, replication, sensitivity analysis, and external checks.

  • Purpose and applications: Simulation supports comparisons, cascading-consequence analysis, sensitivity analysis, and defensive-configuration studies when live experimentation is disruptive, dangerous, or too limited in scale.Applications include stochastic Petri nets for grid design, cyber-physical simulations, system dynamics, and dependency models checked against expert interviews.
  • Protocol: Simulation answers “what changes if” well but answers “what is true” badly, so the research question must be framed as a configuration comparison or parameter sensitivity.The protocol requires the model’s state variables, transitions, stochastic assumptions, and time semantics to be specified.
  • Protocol: Every parameter must be identified as measured, estimated, or assumed, because sensitivity analysis is meaningful only over parameters whose status and plausible range are explicit.Unlabelled parameter tables allow readers to mistake assumptions for measurements.
  • Validity and reporting: Verification must establish that the implementation computes the model, while validation tests whether the model matches reality; both are required.The protocol also requires run length and replication counts based on output variance, with intervals reported over replications rather than a single run.
  • Validity and scope: Simulation predictions have limited transferability beyond the modelled infrastructure and scenarios because model fidelity may not reproduce real organisational responses under time pressure.The reviewed studies used grids, transport, chemical processes, and a hospital as proxies rather than instances of a general enterprise population.

C. Metrics, and why accuracy is the wrong one

Detection evaluation must begin with the operational task, data provenance, and deployment question rather than with a preferred metric or benchmark. Accuracy can be misleading under severe imbalance, while temporal validity, realistic base rates, variance, costs, and transparent preprocessing determine whether comparisons are interpretable.

  • Metrics: Accuracy is dominated by benign traffic under severe class imbalance, so an always-benign classifier can score well without useful detection capability.The relevant operator-facing quantity remains low unless the false-alarm rate is driven extraordinarily low, even with excellent true-positive performance.
  • Metrics: Precision–recall curves are more informative than ROC curves when deployment ranking occurs under heavy imbalance, because ROC false-positive rates are dominated by the large negative class.The metric choice should match the operating question rather than default benchmark practice.
  • Data and validity: A sound protocol audits benchmark provenance and labels, splits before preprocessing, resamples only within training folds, and matches random or temporal splitting to the deployment question.Random k-fold asks whether a model can separate, whereas temporal splitting asks whether it will work in a future period.
  • Baseline protocol: Detection studies must define the operational task, latency, and false-positive versus false-negative costs before selecting data or metrics.The consequence asymmetry determines the metric and operating point and cannot be decided after evaluation.
  • Comparison and reporting: Comparative claims require repeated seeds, an appropriate statistical test, effect sizes, full confusion-matrix reporting, realistic false-alarm rates, and training and inference costs.For randomized procedures, the recommended practice includes enough repetitions, a non-parametric test such as Mann–Whitney U, and a standardized effect size such as Vargha–Delaney A12.

XI. FAMILY H: EVALUATING LANGUAGE-MODEL AND AGENTIC SYSTEMS

Language-model and agentic evaluations introduce validity threats from benchmark contamination, judge-based scoring, adaptive environments, and variable sampling or attempt budgets. The protocol therefore treats contamination checks, calibration, environment specification, pass-at-k, and cost reporting as prerequisites for defensible claims.

  • Validity threats: Language-model evaluations must address unknown training exposure, free-text judging, environmental dependence, and attempt budgets because each creates a distinct validity threat.These threats are especially relevant when systems act agentically rather than merely producing fixed outputs.
  • Contamination: Benchmark contamination must be checked first; post-cutoff or perturbed holdouts, model version and date reporting, and treating public scores as upper bounds are recommended mitigations.A strong score on a benchmark included in training may measure recall rather than capability.
  • Judging: Judge reliability does not establish validity, so judge calibration against human labels is required when model outputs are scored by another model.A judge may agree with itself consistently without tracking the quality being claimed.
  • Agent evaluation: Agent benchmarks must specify the model, version, access date, sampling temperature, attempts per task, and environment, while reporting pass-at-k rather than presenting best-of-n as one attempt.Because success is path-dependent, the environment is part of the measurement.
  • Evidence strength: Operational-capability claims for agentic security systems remain limited because the evaluation methodology is younger than the systems it measures.The paper assigns moderate confidence to contamination control, judge calibration, and environment specification as necessary methodological safeguards.

XII. FAMILY I: MEASUREMENT, TELEMETRY AND DETECTION EFFICACY

Measurement and telemetry research asks what adversary behaviour is observable before asking what a detector can identify. The protocol establishes independent ground truth, measures coverage before efficacy, reports configuration and denominators, and separates logging, detection, alerting, and actioning.

  • Motivation and evidence: Telemetry coverage is a first-order determinant of detection outcomes and is under-measured relative to model performance.A detector cannot find behaviour that the telemetry does not record, so apparent laboratory–production gaps may arise before model choice.
  • Instruments: Honeypot findings are conditional on configuration, so studies must state the emulated device set, exposure period, and network placement before interpreting interaction frequencies.Configuration affects both honeypot attractiveness and the behaviour captured.
  • Protocol: The observable should be defined as adversary behaviour rather than a product feature, with ground truth established independently through emulation, injection, or labelled incidents.Independent ground truth prevents measurements from becoming circular.
  • Protocol: Coverage must be measured before efficacy: first determine what fraction of behaviour reaches telemetry, then what fraction of that observable behaviour is detected.The two quantities have different denominators, so reporting efficacy alone cannot distinguish a weak detector from an absent log source.
  • Configuration and outcomes: Telemetry studies should report agent versions, rule sets, enabled log sources, retention, and the separate stages of logged, detected, alerted, and actioned.A behaviour may be logged, detected, and alerted yet still reach no analyst.

B. The design space

The design space spans reproducible testbeds, emulation, penetration testing, controlled behavioural experiments, and training or range-based evaluation. Across these methods, validity depends on explicit environmental claims, containment, pre-specified protocols, independent instrumentation, and outcome measures tied to the intended decision.

  • Testbeds and ranges: Testbeds organise scenarios, functions, tools, and architectures, while reproducibility, federation, and automated scenario generation address configuration, scale, and preparation costs.Reproducible environments preserve generating configurations with datasets; federated testbeds add topology and heterogeneity; automated generation reduces manual scenario-building effort.
  • Emulation: An emulation plan converts red-team behaviours into repeatable measurements by mapping actions to a technique taxonomy and scoring defensive-stack outputs.The plan is written before execution and run identically across configurations, making endpoint results comparable and turning exercises into data rather than anecdotes.
  • Penetration testing: Penetration-testing research combines scanning, testing, severity classification, and actionable risk reporting, but ambiguous scoping and delivery can undermine comparability.Mixed-method audits illustrate a defender-oriented reporting form, while stakeholder evidence identifies comparability problems in assessment practice.
  • Reproducibility and validity: Testbed protocols should specify environments as code, state falsifiable fidelity boundaries, prevent live-network egress, instrument independently, and publish configuration, plans, and data together.A testbed result transfers imperfectly to production when fidelity claims are not falsifiable; confidence is moderate for controlled experimentation but limited for production transfer.
  • Controlled behavioural studies: Behavioural interventions require randomisation, a contamination-resistant unit, a named comparator, pre-registration before assignment, ethics approval for deception, and intention-to-treat analysis.CONSORT reporting exposes allocation, concealment, assigned-versus-analysed units, follow-up, and effect sizes with intervals.
  • Training and education: Training and range-based education can improve engagement and short-term performance indicators, but long-term effectiveness, organisational change, and scalability remain inconsistently evaluated.The evidence includes adaptive progression, interactive exercises, and a randomised redesign study, while the review assigns limited confidence to sustained outcomes.

D. Awareness programme evaluation

Awareness evaluation must measure behaviour and operational outcomes rather than training completion, using designs that prevent contamination and pre-specify outcomes. The evidence supports moderate short-term gains from interactive approaches, but sustained organisational effects remain uncertain.

  • What to measure: Training completion measures policy compliance, not awareness; systematic evaluation should instead address impact, sustainability, accessibility, and monitoring.The framework adapts four indicators from the European Literacy Policy Network, while small and medium-sized enterprises remain evidence-poor contexts.
  • Ecological validity: Awareness research should distinguish how risk is represented to participants because that design choice affects whether behaviour reflects real stakes.Usable-security research reviews specifically examine risk representation as an empirical design decision.
  • Evidence strength: Interactive exercises and adaptive progression have moderate support for improving engagement and short-term performance indicators, but long-term effectiveness is supported only with limited confidence.Sustained retention, organisational behaviour change, and scalability are not consistently evaluated in the reviewed corpus.
  • Capability evaluation: Operational research evaluates capabilities rather than components, so maturity claims should pair a maturity instrument with an independent outcome measure.A maturity level that correlates with nothing observable describes documentation rather than capability.
  • Experimental design: A behavioural intervention should state its target behaviour, randomise at a contamination-resistant unit, name the comparator, and pre-register outcomes, follow-up, and analysis.Teams or departments are often preferable to individuals because colleagues communicate across conditions.
  • Ethics and reporting: Phishing studies require ethics approval covering employee deception, debriefing, and outcome measurement based on behaviour rather than completion.CONSORT reporting should reconcile assigned and analysed units and report effect sizes with intervals.

B. Prioritising organisational change under disagreement

Organisational-change research must treat prioritisation as a contingent preference structure and validate proposed processes or maturity models against independent outcomes. The review also stresses explicit validity questions, reporting standards, reproducibility, ethics, and artifact safety.

  • Prioritisation: A fuzzy Delphi panel of 29 professionals elicited zero-trust adoption criteria and CRITIC weighted them, treating expert judgements as imprecise rather than point estimates.The resulting ranking is an opinion-based preference structure, not automatically a stable organisational finding.
  • Prioritisation: Multi-criteria rankings should include sensitivity analysis because changing weights can invert a contingent ordering.Reporting a ranking without testing stability presents a preference snapshot as a finding.
  • Validity: Validity review should ask whether measures represent the construct, whether effects have alternative causes, which population the result extends to, and whether analysis supports the inference.Recurring threats include benchmark-construct gaps, leakage, uncontrolled confounds, uncorrected multiple comparisons, single runs, and missing intervals.
  • Process and maturity models: Process-model proposals should reconcile existing models, map each step to an evidential, operational, or regulatory requirement, model dependent tools, use real cases, and pair the model with an independent outcome.The digital-forensics literature frames harmonisation and tool-error enumeration as prerequisites for validated processes.
  • Reporting: Most usable reporting standards come from medicine, epidemiology, and social sciences, while five reviewed methodology families lacked a located reporting standard.The five were threat modeling, simulation, language-model evaluation, testbeds, and design science.
  • Reproducibility and release: Artifact evaluation must add safety review: among 509 artifacts, 41.60% of prevalent findings were judged potential security concerns under practical usage.Dependencies, network behaviour, credentials, keys, and permissive defaults should be addressed before release.
  • Ethics: Ethical reporting should identify the reasoning framework behind decisions, because approval records committee agreement without documenting the grounds for the decision.The review connects this requirement to consequentialist, deontological, and other ethical frameworks applied to security dilemmas.
  • Evidence traceability: Every claim should trace to a specific artifact, table, run, transcript, or log, with wording calibrated to the strength and verification status of that evidence.A claim-to-evidence table is recommended as a parallel drafting instrument.

XVII. READING THE CONTRADICTIONS

Intrusion-detection studies report incompatible rankings despite careful individual designs, so benchmark accuracy cannot be treated as a stable estimate of enterprise performance. The review attributes the observed inversions to joint evaluation-design variation, while calling for harmonised, multi-dataset, deployment-aware protocols and triangulation.

  • The contradiction: Four careful comparative studies rank intrusion-detection algorithms differently, showing that no stable ordering emerges across the reviewed evaluations.The studies variously favour ensembles, deep feed-forward networks, recurrent networks, or random forests under different designs.
  • The contradiction: 98% accuracy was reported for convolutional and recurrent models, versus 99.9% for random forest, illustrating incompatible benchmark comparisons.These values come from one later comparison and should not be read as a universal algorithm ranking.
  • The resolution: Dataset composition, features, preprocessing, network context, class balance, architecture, and tuning budgets jointly vary enough to account for the observed margins.The review does not isolate any single factor and does not resolve the contradiction by selecting a winner.
  • The resolution: Headline benchmark accuracy is not a stable estimate of enterprise performance, and reporting it as though it were overstates the result.The supported claim concerns joint design variation, not a demonstrated effect of one isolated design factor.
  • Deployment uncertainty: A hybrid ensemble reached accuracy, precision, recall, and F1 approaching 100% under SMOTE, weighted soft voting, five-fold cross-validation, and an independent benchmark.The study still reported computational overhead, latency, overfitting risk, and no explicit unseen-attack evaluation.
  • Research implications: Evaluation should use multiple datasets, temporal splits for future-performance claims, realistic-base-rate false alarms, simple baselines, and computational-cost reporting.These design choices address benchmark variation and deployment relevance rather than merely increasing headline accuracy.
  • Research gaps: The field’s strongest unresolved gap is evaluation in continuously operating enterprises with changing attack distributions, incomplete telemetry, multiple organisations, and unseen-attack testing.The review also calls for harmonised preprocessing, imbalance treatment, hyperparameter reporting, cost measurement, and train-test protocols.
  • Synthesis: Enterprise studies should combine technical, organisational, and longitudinal evidence because communication, decision load, human factors, intelligence integration, and temporal behaviour shape security outcomes.The conclusion frames methodological combination and triangulation as necessary to assess changing enterprise conditions.

APPENDIX A CONSOLIDATED REPORTING CHECKLISTS

Table IV consolidates the reporting checklist items most frequently absent from the reviewed corpus, organised by methodology family.

  • The checklist is intended for both writing and reviewing enterprise cybersecurity research.
Loading 2608.24850v1…