Source-linked AI summary
Infrastructure Resilience Curves: Performance Measures and Summary Metrics
Craig Poulin, Michael Kane
TL;DR
Resilience-curve analysis lacks clear guidance for selecting performance measures and summary metrics, even though those choices shape infrastructure assessments. Through a critical review of 273 publications, the paper develops taxonomies and a framework for examining assumptions. It recommends broader use of productivity measures, attention to milestones and endogenous targets or thresholds, and caution when summarizing ensembles of curves.
Problem
Resilience literature lacks clear guidance for selecting performance measures and summary metrics despite their potential to yield different recommendations.
Method
The manuscript critically reviews 273 publications to develop taxonomies of resilience-curve performance measures and summary metrics and synthesize selection recommendations.
Results
The review defines taxonomies for performance measures and summary metrics and identifies summary-function and metric-selection patterns across the literature.
Takeaways & Limitations
Resilience-curve analyses should align measures and metrics with stakeholder goals and deliberately consider curve milestones when defining summaries.
Takeaways & Limitations
For ensembles of curves, the authors recommend reporting descriptive statistics for each curve’s metric rather than deriving metrics from an expected trajectory.
Abstract
from arXiv · showhide
Resilience curves are used to communicate quantitative and qualitative aspects of system behavior and resilience to stakeholders of critical infrastructure. Generally, these curves illustrate the evolution of system performance before, during, and after a disruption. As simple as these curves may appear, the literature contains underexplored nuance when defining "performance" and comparing curves with summary metrics. Through a critical review of 273 publications, this manuscript aims to define a common vocabulary for practitioners and researchers that will improve the use of resilience curves as a tool for assessing and designing resilient infrastructure. This vocabulary includes a taxonomy of resilience curve performance measures as well as a taxonomy of summary metrics. In addition, this review synthesizes a framework for examining assumptions of resilience analysis that are often implicit or unexamined in the practice and literature. From this vocabulary and framework comes recommendations including broader adoption of productivity measures; additional research on endogenous performance targets and thresholds; deliberate consideration of curve milestones when defining summary metrics; and cautionary fundamental flaws that may arise when condensing an ensemble of resilience curves into an "expected" trajectory.
1 Introduction
Resilience curves represent system performance through disruption and recovery, but selecting performance measures and summary metrics remains context-dependent and underexplored. This survey develops broadly applicable taxonomies and recommendations for using curves across critical infrastructure analysis.
- Resilience curves: Resilience curves show how a selected performance measure changes over time before, during, and after disruption.They may represent empirical or simulated behavior and can include cascading failure, degraded response, non-increasing recovery, or incomplete restoration.
- Resilience curves: A performance measure maps system states to a scalar over time, whereas a summary metric maps an entire resilience curve to a scalar.Both should reflect stakeholder interests and goals.
- Analytical choices: Different measures and metrics can yield significantly different recommendations, and a single metric cannot capture all characteristics of a resilience curve.Examples include divergent water-system measures, threshold-based versus cumulative metrics, and dissimilar curves with equal summary metrics.
- Analytical choices: The literature lacks clear, general guidance for selecting measures according to stakeholder values and goals, partly because selection imposes non-trivial analytical and data-collection burdens.Measure applicability and its consequences for analysis effort are also underexplored.
- Survey contribution: This manuscript critically surveys 273 publications to define taxonomies of performance measures and summary metrics and synthesize best practices for their selection and communication.Its scope is resilience-curve measures and metrics rather than general infrastructure resilience or non-curve metrics.
2 Literature Survey
The manuscript presents a critical review of critical infrastructure resilience-curve literature, using broad searches and inclusion criteria to synthesize trends and opportunities. The survey covers 273 publications spanning infrastructure types, disruption types, and publication formats.
- Review approach: The review synthesizes resilience-curve measures and metrics through a critical review intended to highlight trends and opportunities.Its flexible synthesis is broader but more subjective than a systematic review.
- Publication selection: Filtering retained 184 publications, representing 13.3% of the 1,384 publications accessible through Northeastern University licenses.An additional 89 publications were added from references during review.
- Survey scope: The final survey scope comprised 273 publications across journal articles, conference papers, book sections, magazine articles, and one report.The collection included 214 journal articles from 101 journals and 45 conference papers.
- Survey scope: The surveyed literature covered energy, transportation, water, financial, information, healthcare, supply-chain, and coupled infrastructure systems.Disruptions included natural and manmade events lasting from minutes to days.
3 Performance Measures
The manuscript organizes performance measures into availability, productivity, and quality categories, while also addressing ensemble measures and normalization. These categories differ in what aspect of infrastructure service they represent and in their modeling requirements.
- Taxonomy: The performance-measure taxonomy defines availability, productivity, and quality categories, alongside ensemble measures and three normalization schemes.The categories are summarized in Table 2.
- Availability: Availability measures describe infrastructure capacity or functionality and are generally straightforward to model when service-demand dynamics can be excluded.They suit stakeholders focused on the infrastructure system rather than end-use service.
- Productivity: Productivity measures describe the quantity or flow of service, requiring consideration of supply, capacity, and demand.They may be most appropriate when service demand changes with infrastructure condition and for end-use stakeholders.
- Quality: Quality measures describe the character of provided service and can reveal tradeoffs between steady-state and disrupted performance.They require modeling how supply and demand affect service quality, increasing analytical complexity.
- Ensemble measures: Ensemble measures require stakeholder-informed weighting, but identifying, interpreting, and balancing divergent or contradictory measures remains difficult.Investment alternatives varied widely with weighting across five perspectives.
3.3 Performance Normalization
The manuscript distinguishes static, exogenous, and endogenous normalization according to how the reference value changes and responds to scenario dynamics. It shows that normalization choices affect comparability, interpretation, and analytical recommendations.
- Interpretation: Normalized values can obscure important context, so presenting both actual and normalized performance resolves this disadvantage.The issue is especially relevant when systems differ substantially in scale or population.
- Normalization schemes: Normalization functions are underexplored, and some publications omit or obscure their denominators despite normalization being used to compare systems and scenarios.Changes in the denominator can be as impactful as changes in 𝒫𝒫(t).
- Normalization schemes: The manuscript defines static, exogenous, and endogenous normalization schemes for translating unnormalized performance 𝒫𝒫(t) into normalized performance 𝑝𝑝(t).Static uses ℛ0, exogenous uses a time-varying nonresponsive baseline ℛ(t), and endogenous uses a scenario-responsive target ℛ(t).
- Static normalization: Static normalization was used extensively for availability measures, where ℛ0 generally represented nominal or upper-bound performance and emphasized restoration to full function.Static normalization also appeared for productivity and quality measures, but with less straightforward interpretations.
- Exogenous normalization: Exogenous normalization accommodates time-varying demand but does not model interactions between the scenario and service demand.Timing of system failures can affect summary metrics when demand varies hourly.
- Endogenous normalization: Endogenous normalization uses a time-varying performance target affected by the hazard or system response and is applicable when full recovery is not possible or feasible.Productivity measures are identified as ideal candidates, although modeling the target may be as intensive as modeling performance.
4 Summary Metrics
The manuscript classifies summary metrics by what they extract from resilience curves and emphasizes that metric definitions depend on units, milestones, reference values, and control intervals. It finds no consensus on a best metric and warns that curve structure and assumptions must be made explicit.
- Metric taxonomy: Summary metrics map actual or normalized resilience curves to scalar values for comparing system behavior across scenarios and configurations.The taxonomy defines six categories and provides examples and best practices.
- Metric taxonomy: The survey found no consensus on a best summary metric; metrics were instead linked to desired attributes.Only one publication directly compared metrics, and threshold- and integral-based metrics produced different recommendations.
- Formulation considerations: Metric formulation should specify performance units, milestones or criteria, reference values, duration, and other category-specific design choices.Metrics may use time, performance × time, binary values, or generally unitless indices.
- Milestones and control times: Metrics are commonly defined over a control interval [t0, tc] and with curve milestones that delineate transitions between phases.Milestones may be absent, additional, ambiguous, or unclear for trajectories with temporary performance increases.
- Milestones and control times: Resilience analysis should clearly define milestones and validate those definitions across the range of considered trajectories.This is necessary because common curve forms do not guarantee essential, comprehensive, or unambiguous milestones.
- Milestones and control times: Control-time selection can use expected recovery, maximum recovery, stakeholder-defined duration, or a hazard-lifecycle horizon, each affecting curve truncation or interpretation.Some choices are undefined for systems that do not fully recover, while lifecycle analyses may require discount rates and adaptation considerations.
4.2 Magnitude-based Metrics
Magnitude-based metrics quantify performance at selected milestones or times, including residual, depth-of-impact, and restored-performance measures. Their interpretation depends on clearly defined milestones, reference values, and normalization schemes.
- Magnitude-based metrics quantify performance at a specific milestone or point in time.
- Residual performance: Residual performance metrics describe system performance following disruption, often relative to a critical threshold or average disruption-period performance.
- Depth of impact: Depth-of-impact metrics complement residual performance and are associated most often with robustness, alongside absorptive capacity, survivability, and vulnerability.
- Residual performance: Residual performance requires a clearly defined milestone, such as post-hazard, post-cascading-failure, or minimum performance.
- Restored performance: Restored performance metrics quantify performance after recovery efforts, including partial recovery and permanent outages.
- Residual and restored-performance calculations vary with the reference and normalization scheme, while endogenous targets can decouple minimum performance from physical degradation.
4.3 Duration-based Metrics
Duration-based metrics quantify time between disruption and recovery milestones, but comparisons depend on consistent milestone definitions. Integral-based metrics combine performance and time, with normalization and control-duration choices shaping interpretation.
- Duration-based metrics: Duration-based metrics quantify time between milestones, covering disruption, absorb, endure, and recovery phases.
- Milestones: Recovery-duration milestones are often ambiguous because authors may conflate cascading-failure ends with recovery onset or use resilience-triangle conventions.
- Milestones: Recovery time is undefined when a system never fully recovers; specified thresholds, such as restoration to 95% of ridership, provide an alternative.
- Duration-based metrics: Most duration metrics prefer shorter disruptions or faster recoveries, although some forms prefer higher values such as uptime or speed-recovery factors.
- Integral-based metrics: Integral-based metrics incorporate both time and performance through unnormalized or time-, performance-, or jointly normalized forms.
- Integral-based metrics: The manuscript uses “cumulative performance” rather than “resilience” because summary metrics do not fully describe a system and the term resilience has varied uses.
- Integral-based metrics: Both performance and time were most commonly normalized, producing a unitless metric, while fixed control intervals can obscure curve nuance or confuse comparisons.
- Integral-based metrics: Integral calculations commonly assume equal value across performance-time units, although stakeholder weighting and nonlinear value functions offer alternatives.
4.5 Rate-based Metrics
Rate-based metrics quantify how performance changes over time, typically through derivatives or approximations of failure and recovery phases. They require clearly specified milestones and have direction-dependent preferred values.
- Rate-based metrics quantify how system performance changes over time, commonly using derivatives or linear approximations.
- Failure rate: Failure-rate metrics describe performance loss during failure or combined endure-and-recovery phases, with lower-magnitude negative derivatives generally preferred.
- Ensembles: The distinction between derivative sign and magnitude matters when rate metrics are used in an ensemble.
- Recovery rate: Recovery-rate metrics quantify restorative capability across recovery milestones, using derivatives, linear approximations, arctan extensions, or exponential-decay fits.
- Recovery rate: Higher recovery-rate values are preferred because they represent steeper recovery.
- Milestones: Rate-based metrics require clearly specified milestones, and poorly defined milestones can make implementation challenging.
4.6 Threshold-based Metrics
Threshold-based metrics represent nonlinear transitions in performance interpretation through critical-performance and recovery-time thresholds. They either assess threshold adherence categorically or modify other metric forms.
- Thresholds represent nonlinear transitions in quantitative interpretation and may mark discontinuities for summary metrics.
- Threshold types: Critical-performance thresholds identify unacceptable performance, while recovery-time thresholds specify when objectives should be achieved.
- Threshold adherence: Threshold-adherence metrics provide categorical system assessments based on maintaining performance above a critical threshold, recovering within a time threshold, or satisfying both.
- Threshold modification: Threshold-modified metrics adjust forms from another metric category using performance or recovery thresholds.
- Examples: Examples include residual capacity as residual performance minus a critical threshold, recovery-time adjustments, and brittleness as impact below a critical threshold.
4.7 Ensemble Summary Metrics
Ensemble summary metrics consolidate multiple metrics, performance measures, or scenarios into single values, but this can obscure important distinctions. The review especially cautions that metrics derived from expected trajectories may misrepresent scenario ensembles.
- Ensemble metric categories: Ensemble summary metrics combine metric, measure, or scenario ensembles into a single value for optimization or concise communication.The three categories consolidate multiple metrics for one curve, multiple measures, or possible system behaviors across scenarios.
- Metric ensembles: Metric ensembles combine distinct metrics from the same curve, often using unitless sums or products and sometimes incorporating financial considerations.Constituent metrics may have different units; financial conversion can also include non-curve considerations such as recovery costs.
- Measure ensembles: Measure ensembles consolidate multiple performance measures, typically by summarizing distinct curves with a common integral-based metric.Normalized metrics enable combinations when measures have dissimilar units, including geometric means and products.
- Measure ensembles: Weighting schemes can encode stakeholder preferences, subsystem interdependence, connectivity, time variation, or relationships estimated from historical data.The review identifies these as underexplored opportunities for ensemble metrics across performance measures.
- Scenario ensembles: Scenario ensembles summarize metric values across enumerated scenarios or simulations, but distributions were infrequently presented compared with single values.Examples include histograms, distribution functions, percentiles, and probabilities of remaining within thresholds.
- Expected trajectories: Expected trajectories may not correspond to any constituent trajectory, and magnitude-, duration-, rate-, and threshold-based metrics derived from them can differ from every scenario.Integral-based metrics may reflect expected constituent metrics in the illustrated example, whereas other metric categories do not necessarily do so.
4.8 Summary Functions
Summary functions evaluate resilience-curve behavior at particular times rather than necessarily reducing an entire curve to one scalar. Recovery ratio is the most common example and measures improvement relative to minimum performance.
- Summary functions: Twenty-one publications implemented functions that do not map an entire resilience curve to a scalar value.These functions can be evaluated at any point within the scenario and therefore are not summary metrics in the manuscript’s terminology.
- Summary functions: Examples include local resilience as the derivative of a resilience curve and space-time dynamic resilience as normalized cumulative performance since disruption.The manuscript avoids labeling these functions as variations of “resilience.”
- Recovery ratio: Recovery ratio quantifies performance improvement at any time relative to the curve’s minimum performance.It was the most common summary function across the reviewed infrastructure sectors.
- Recovery ratio: Recovery ratio commonly presumes availability measures and a fixed reference ℛ0, although extensions can incorporate performance targets ℛ(t).The function may consequently be used to represent restoration of maximum potential network flow.
- Interpretation: A recovery ratio of 0 at disruption time does not mean zero resilience, and a value of 1 indicates recovery rather than full resilience.The metric can support restoration sequencing and component-importance estimation.
5 Discussion
The discussion argues that resilience analyses should align performance measures, normalization, summary metrics, communication, and stakeholder goals with system realities. It emphasizes productivity measures, dynamic targets, descriptive ensemble statistics, and caution when interpreting expected trajectories.
- Discussion: Incorrectly selecting or implementing resilience-curve tools can yield incorrect results, so analyses should address measures, metrics, communication, and practice.The discussion identifies these four aspects as requiring careful consideration before use.
- Performance measures: Quality measures can represent service character and support trade-off analysis, but may be inappropriate when productivity goals cannot be met or disruptions do not affect quality.They can require modeling service utilization and other non-infrastructure considerations.
- Normalization: Availability measures may overvalue excess capacity and undervalue dynamic, real-time resilience when static normalization is adopted inappropriately.Availability measures are simpler in some settings but should reflect stakeholder goals and system realities.
- Performance measures: Productivity measures incorporate infrastructure service supply and demand and can expose additional resilience-improving intervention options.They are especially relevant to downstream stakeholders concerned with service provision.
- Performance targets: Productivity is understood relative to demand, so demand should generally be modeled as dynamic unless constant or exogenous baselines are justified.Dynamic targets can relate infrastructure states to stakeholder goals and behaviors.
- Communication of results: Expected trajectories can mislead stakeholders because metrics derived from them may be objectively incorrect for an ensemble of curves.Descriptive statistics for each curve’s metric and likelihoods of specific scenarios are proposed alternatives.
- Practice of resilience analysis: Stakeholder goals and values may not fully align, making stakeholder identification and engagement relevant to defining performance.The review also identifies empirical analysis and calibration as areas warranting additional research.
6 Conclusion
The conclusion presents a vocabulary for resilience-curve analysis based on a critical review of 273 publications. It defines performance-measure and summary-metric taxonomies and recommends careful milestone, normalization, and ensemble interpretation.
- Contribution: 273 publications informed taxonomies of resilience-curve performance measures and summary metrics, plus recommendations for selecting and applying them.The review is intended to support future critical infrastructure analysis.
- Performance measures: Three performance-measure categories are defined in increasing modeling complexity: availability, productivity, and quality.Availability concerns capacity or aggregated function, productivity concerns service quantity, and quality concerns service character.
- Normalization: Normalization is classified as static, exogenous, or endogenous, with static normalization generally associated with availability measures and full recovery goals.Multiple measures or spatial variations may require ensembles.
- Summary metrics: Six summary-metric categories are defined: magnitude, duration, integral, rate, threshold, and ensembles.Summary metrics distill curves to single values to facilitate comparison.
- Metric interpretation: Summary metrics should be tied to clearly defined curve milestones because assumptions such as instantaneous loss or non-decreasing recovery affect their effectiveness.The conclusion specifically recommends careful selection and specification of metrics.
- Ensemble analysis: For curve ensembles, analyses should report descriptive statistics for each curve’s metric rather than derive metrics from an expected trajectory.Expected-trajectory representations can mislead stakeholders and produce objectively incorrect metrics.
- Future research: Future research should examine social and technical interactions, stakeholder engagement, and validation of stakeholder-defined performance interpretations.The conclusion frames resilience curves as one tool within broader resilience analysis.