Source-linked AI summary

Operational resilience: concepts, design and analysis

Alexander A. Ganin, Emanuele Massaro, Alexander Gutfraind, Nicolas Steen, Jeffrey M. Keisler, Alexander Kott, Rami Mangoubi, Igor Linkov

arXiv:1508.01230v1physics.soc-phphysics.data-an

TL;DR

Traditional risk-based approaches face difficulty addressing widely unknown challenges in interconnected critical infrastructure. The paper develops a time-based resilience framework and shows that network parameters can be traded off to obtain desired resilience and robustness levels.

  • Problem

    Traditional risk-based approaches are criticized for difficulty addressing widely unknown challenges in interconnected critical infrastructure.

  • Method

    The paper implements the NAS definition of resilience as a time-dependent function using critical functionality measures across directed acyclic graphs and interdependent networks.

  • Results

    Network parameters can be traded off to obtain desired overall resilience and robustness levels.

  • Takeaways & Limitations

    The framework supports evaluating resilience across time while examining design tradeoffs in directed acyclic graphs and interdependent networks.

  • Takeaways & Limitations

    The directed acyclic graph model assumes links run only from higher levels to lower levels.

Abstract

from arXiv · show

Building resilience into today's complex infrastructures is critical to the daily functioning of society and its ability to withstand and recover from natural disasters, epidemics, and cyber-threats. This study proposes quantitative measures that implement the definition of engineering resilience advanced by the National Academy of Sciences. The approach is applicable across physical, information, and social domains. It evaluates the critical functionality, defined as a performance function of time set by the stakeholders. Critical functionality is a source of valuable information, such as the integrated system resilience over a time interval, and its robustness. The paper demonstrates the formulation on two classes of models: 1) multi-level directed acyclic graphs, and 2) interdependent coupled networks. For both models synthetic case studies are used to explore trends. For the first class, the approach is also applied to the Linux operating system. Results indicate that desired resilience and robustness levels are achievable by trading off different design parameters, such as redundancy, node recovery time, and backup supply available. The nonlinear relationship between network parameters and resilience levels confirms the utility of the proposed approach, which is of benefit to analysts and designers of complex systems and networks.

Introduction

The paper addresses limitations of threat-specific risk assessment by proposing a quantitative, time-dependent measure of engineering resilience based on stakeholder-defined critical functionality. It applies the approach to directed acyclic graphs and interdependent coupled networks, including an Ubuntu code-system approximation.

  • Introduction: The approach responds to risk-based methods’ difficulty addressing widely unknown and uncertain threats, whose realistic scenarios and hardening investments can be difficult to justify.
  • Introduction: The paper proposes a methodology that quantifies NAS-defined engineering resilience using stakeholder-defined critical functionality and its temporal recovery profile.
  • Introduction: Critical functionality can measure integrated resilience and support additional performance metrics, including robustness, through indicators such as functioning-node percentage and flow-to-capacity ratio.
  • Introduction: The methodology focuses on multi-level directed acyclic graphs and interdependent coupled networks, with an Ubuntu 12.04 code system approximated as a directed acyclic graph.
  • Introduction: Because analytical results are generally intractable, the approach is simulation based, while a simple special case yields analytical insight into redundancy and system resilience.

Resilience: an analytical definition

The paper defines resilience as a quantitative measure of critical functionality over time under specified adverse events, integrating network structure and temporal evolution. The framework evaluates resilience through simulation and applies it to hierarchical DAGs and interdependent coupled networks.

  • Definition: Critical functionality K maps system states or parameters to a value between 0 and 1, using node or link importance and activity or functionality.The activity term π_i(t; C) can represent the degree to which an element remains active under attack or its probability of being fully functional.
  • Definition: Resilience R is a composite function of network topology, temporal evolution, critical functionality, and a specified class of adverse events.The evaluation interval [0, T_C] is set by stakeholders or estimated from the mean time between adverse events.
  • Definition: The resilience measure captures normalized performance before, during, and after an attack, including preparation, absorption, recovery, and adaptation.The formulation can use continuous or discrete time and is commonly normalized to K_nominal(t) = 1.
  • Evaluation method: Because closed-form evaluation over all adverse events is infeasible, the approach simulates possible system evolutions and averages critical functionality at each time step.Resilience is then calculated over the control interval from the simulated functionality trajectories.
  • Adverse events: Unlike probability-weighted risk-based approaches, the framework defines resilience for damage caused by a specified adverse event or event class regardless of its occurrence probability.This follows the NAS framing of resilience as the ability to plan, absorb, recover, and adapt.
  • Models and applications: The framework is demonstrated with hierarchical multi-level DAGs and interdependent coupled networks, including synthetic supply-demand graphs and the Ubuntu 12.04 software network.The DAG analysis examines redundancy, switching, and recovery-time tradeoffs, while the coupled-network model represents mutual interdependence between networks A and B.

Results · Model 1 – Directed acyclic graphs. Synthetic graphs

Results across directed acyclic and interdependent network models show that resilience depends nonlinearly on damage, recovery, redundancy, switching, and backup resources. Synthetic and Linux-network analyses identify parameter tradeoffs and failure conditions that determine critical functionality and recovery.

  • Model 1 – Directed acyclic graphs. Synthetic graphs: Short recovery times can yield resilience values close to 1 even when ps is as small as 0.05, and recovery time affects resilience more strongly than increased redundancy.This result holds for TR < 0.1TC, with redundancy held constant at pm = 0.01.
  • Model 1 – Directed acyclic graphs. Linux software network: In the Linux package network, random attacks produce R = 0.99975 and M = 0.999, significantly exceeding guided attacks because failed nodes are generally less important.Guided-attack damage depends on which packages are targeted.
  • Model 2 – Interdependent networks. Synthetic graphs: With insufficient backup agents, interdependent ER networks cannot return to recovery and critical functionality oscillates between 0 and about 0.5.Backup supply to 0.4N nodes in network A allows the system to reach a stable state after backup removal and further recovery.
  • Model 2 – Interdependent networks. Synthetic graphs: In scale-free networks, recovery is more dispersed and depends strongly on whether important hubs are affected during cascading failure.Unaffected hubs correspond to relatively small damage, whereas affected hubs can cause a large critical-functionality drop and no full recovery within the control period.
  • Model 2 – Interdependent networks. Synthetic graphs: When the entire network A is destroyed, recovery cannot start because network A lacks a giant component.The same general tendencies occur in scale-free networks, but their phase-diagram region is more dispersed.

Conclusion

The paper presents a time-dependent approach for implementing the NAS definition of resilience and applying it to design tradeoffs in complex networks. It shows how parameter choices can be evaluated for their effects on resilience and robustness, while identifying adaptation and broader network types as future work.

  • Conclusion: The approach implements the NAS definition of resilience as a function of design tradeoff parameters in directed acyclic graphs and interdependent networks.
  • Conclusion: Evaluating resilience across time rather than as a single quantity enables designers to analyze how parameter choices and design changes affect network resilience and robustness.
  • Conclusion: Trading off network parameters can achieve desired resilience and other performance-measure levels in the demonstrated network classes.
  • Conclusion: Future work will extend the approach to multiplex and other real-life networks and model adaptation after restoration to improve resistance to future adverse events.

Methods

The study models resilience using hierarchical directed acyclic graphs and interdependent coupled networks, incorporating dependency-driven failures, recovery, and backup-supply switching. Analytical approximations describe active-node dynamics under small damage and instantaneous switching assumptions.

  • Hierarchical multi-level DAG: The hierarchical DAG organizes supply–demand links across levels, with links directed from higher to lower levels and no same-level links or cycles.Each node has real supply links and may have virtual backup links from potential suppliers.
  • Hierarchical multi-level DAG: Nodes become inactive through destruction or unresolved dependencies, while eligible disabled nodes can switch virtual links to real links with probability ps.Switching may be instantaneous or delayed, and node deactivation can propagate through dependencies lacking alternative real suppliers.
  • Analytical approximation: The analytical approximation derives active-node and resilience dynamics for small initial damage, instantaneous switching with ps = 1, and recovery after TR time steps.Destroyed nodes are rebuilt after TR unless they remain unsupplied by upper levels.
  • Interdependent coupled networks: In coupled networks, failures propagate through deactivation, fragmentation, and inter-network dependencies until no further nodes can be removed.Recovery uses Nb backup agents to replace unresolved dependencies, activate connected nodes, and propagate recovery between networks until full recovery or the control-time limit.

Author contributions statement

All authors contributed to the study’s concept and model, while specific authors led software development, experimentation, analysis, guidance, and manuscript review. The study was funded by the US Army, with publication permission from USACE and an authorship disclaimer regarding sponsors’ views.

  • All authors developed the concept and model, while A.G. developed the software and conducted the experiment.
  • A.G., E.M., A.Gu., and I.L. analyzed the results, while I.L. conceived the idea and provided overall guidance.
  • N.S., J.K., A.K., and R.M. reviewed the manuscript; the study was funded by the US Army and approved for publication by USACE.
  • The authors stated that the paper’s views are their own and not those of the US Army or other sponsor organizations.

Figure legends

The figures define resilience through time-integrated critical functionality and illustrate how redundancy, switching, repair, and backup supply shape recovery. Applications to synthetic, Linux, and coupled networks show scenario-dependent damage and stochastic restoration outcomes.

  • Resilience and critical functionality: Resilience is evaluated as the integral of critical functionality K over time, providing the central measure used throughout the figures.The critical functionality is explicitly defined as K’s dependency on time.
  • Network generation and adverse events: The adverse-event model defines hierarchical networks, establishes links probabilistically with pm, permits switching with ps, and restores damaged nodes after TR steps.Nodes with redundant links can switch after attack, while initially destroyed nodes recover after the repair period.
  • Synthetic networks: Synthetic results show robustness values of 0.966, 0.787, and 0.453 across scenarios, demonstrating sensitivity to damage and redundancy conditions.The supplied passage also identifies a fourth scenario, but its robustness value is truncated.
  • Linux network: Linux-network attacks differ substantially in damage: guided attacks yield robustness values from 0.982 to 0.129, whereas random attacks yield M = 0.999.Within guided attacks, libc6 and gcc-4.6-base are most damaging, while xauth is less damaging than libstdc++6.
  • Coupled-network recovery: Recovery in coupled networks is stochastic and highly sensitive to backup supply, with success depending mainly on cascade outcomes rather than random backup-node selection.Recovery is more likely when important hubs remain active after cascading failure; phase diagrams show resilience varying with Nb and pdestr.
Loading 1508.01230v1…