Source-linked AI summary

Flood Prediction Using Machine Learning Models: Literature Review

Amir Mosavi, Pinar Ozturk, Kwok-wing Chau

arXiv:1908.02781v1cs.LGstat.ML

TL;DR

Flood prediction involves highly destructive events and ML capabilities that vary across algorithms. This paper reviews and compares ML methods, identifying major trends for improving prediction quality.

  • Problem

    Floods are highly destructive, while the capability of each ML algorithm may vary across applications.

  • Method

    The paper reviews ML applications in flood prediction, compares performance using error and correlation measures, and examines hybrid methods and ML taxonomies.

  • Results

    The review reports four major trends for improving prediction quality: hybridization, data decomposition, algorithm ensembles, and model optimization.

  • Takeaways & Limitations

    Improving flood-prediction quality is associated in the reviewed literature with hybridization, data decomposition, algorithm ensembles, and model optimization.

  • Takeaways & Limitations

    Some ML methods have relatively low accuracy, require repeated parameter tuning, and respond slowly.

Abstract

from arXiv · show

Floods are among the most destructive natural disasters, which are highly complex to model. The research on the advancement of flood prediction models contributed to risk reduction, policy suggestion, minimization of the loss of human life, and reduction the property damage associated with floods. To mimic the complex mathematical expressions of physical processes of floods, during the past two decades, machine learning (ML) methods contributed highly in the advancement of prediction systems providing better performance and cost-effective solutions. Due to the vast benefits and potential of ML, its popularity dramatically increased among hydrologists. Researchers through introducing novel ML methods and hybridizing of the existing ones aim at discovering more accurate and efficient prediction models. The main contribution of this paper is to demonstrate the state of the art of ML models in flood prediction and to give insight into the most suitable models. In this paper, the literature where ML models were benchmarked through a qualitative analysis of robustness, accuracy, effectiveness, and speed are particularly investigated to provide an extensive overview on the various ML algorithms used in the field. The performance comparison of ML models presents an in-depth understanding of the different techniques within the framework of a comprehensive evaluation and discussion. As a result, this paper introduces the most promising prediction methods for both long-term and short-term floods. Furthermore, the major trends in improving the quality of the flood prediction models are investigated. Among them, hybridization, data decomposition, algorithm ensemble, and model optimization are reported as the most effective strategies for the improvement of ML methods.

1. Introduction

Flood prediction is difficult because floods are complex and existing physical and statistical models face data, computation, accuracy, and short-term forecasting constraints. The review motivates machine learning as a data-driven alternative and identifies a gap in comprehensive comparisons across flood applications.

  • Floods cause extensive damage to human life, infrastructure, agriculture, and socioeconomic systems, motivating reliable prediction and risk management.
  • The review is limited to lead-time prediction at identified sites and does not address spatial flood prediction or flood-location estimation.
  • Physical models can represent diverse flooding scenarios but require extensive monitoring data, intensive computation, and specialized hydrological expertise.
  • Physical and statistical approaches have reported limitations in short-term prediction, including insufficient accuracy, systematic errors, and only moderate forecast skill.
  • Machine learning models formulate flood nonlinearity from historical data without requiring explicit knowledge of underlying physical processes.
  • Machine learning offers relatively fast development, low computation cost, minimal inputs, and high performance compared with physical models.
  • The literature reports effective ML algorithms for short- and long-term forecasting, with performance further improved through hybridization and related combined approaches.
  • Despite many individual evaluations, the literature lacks a definite conclusion about which ML models perform best for particular applications and lacks a comprehensive general review.

2. Method and Outline

The survey organizes flood-prediction literature through a search strategy combining flood-resource variables, ML methods, and prediction terminology. It screens the resulting studies using quality measures and reviews them by variables, methods, prediction type, and results.

  • The survey identifies ML methods for flood prediction and classifies applications by flood-resource variables, prediction type, and obtained results.
  • Articles were prioritized for performance evaluation and comparison and assessed using SNIP, CiteScore, SJR, and h-index.
  • Rainfall is important for runoff and flood modeling, but accurate flood prediction also considers other variables such as soil moisture.
  • The search combines flood-resource variables, ML algorithms, and four prediction terms: prediction, estimation, forecast, and analysis.
  • The flood-resource-variable search covered 25 keywords, while the ML-method search used the 25 most popular engineering algorithms as keywords.
  • The search returned 6596 articles, from which 180 original research papers were refined using the survey’s quality measure.
  • The survey covers short-term and long-term flood prediction and presents the state of the art, technical descriptions, and comparative performance analysis.

3. State of the Art of ML Methods in Flood Prediction

Flood prediction workflows use historical hydrometeorological data divided into training and evaluation sets, with ML algorithms applied to model complex flood relationships. Among the reviewed approaches, radar-based inputs and ANN variants were associated with strong prediction performance, although ANNs retain important limitations.

  • Data and workflow: Flood datasets commonly combine historical flood events with rainfall, water-level, streamflow, or sensor observations across hourly, daily, and monthly timescales.Sources include ground gauges, satellites, multisensor systems, and weather radar.
  • Data and workflow: Radar observations can provide higher-resolution and more reliable data than rain gauges, and radar-based rainfall models were reported to provide higher accuracy in general.
  • Data and workflow: The basic ML workflow divides datasets into individual sets and applies training, validation, verification, and testing to construct and evaluate prediction models.Figure 2 represents this basic model-building flow.
  • Core ML methods: Major flood-prediction algorithms include ANNs, neuro-fuzzy systems, ANFIS, SVMs, WNNs, and MLPs.
  • Artificial neural networks: ANNs model complex nonlinear rainfall–flood and river-flow relationships and were reviewed as suitable because of acceptable generalization ability and speed.ANNs derive meaning from historical data rather than catchment physical characteristics.
  • Artificial neural networks: ANN limitations include network-architecture and data-handling challenges, physical-interpretation difficulties, relatively low accuracy, repeated parameter tuning, and slow gradient-based learning.
  • Artificial neural networks: BPNNs were identified as powerful tools for flood time-series prediction, while ELMs were used for short-term streamflow modeling with promising results.

3.2. Multilayer Perceptron (MLP)

MLP is a multilayer feed-forward ANN trained with supervised backpropagation. In flood-modeling assessments, MLPs were reported to offer efficient performance and better generalization, but can be difficult to implement.

  • Definition and structure: MLP is a feed-forward ANN representation that uses supervised backpropagation to train interconnected nodes across multiple layers.
  • Definition and structure: Simplicity, nonlinear activation, and a high number of layers characterize MLPs used in flood prediction and other complex hydrogeological models.
  • Performance and limitation: MLP models were reported to be more efficient and to have better generalization ability than other ANN classes used in flood modeling.
  • Performance and limitation: MLP is generally found to be more difficult to use despite its reported efficiency and generalization advantages.
  • Use in flood prediction: MLP gained popularity among hydrologists and was treated as a separate model class because numerous flood-prediction studies used the MLP designation.

3.3. Adaptive Neuro-Fuzzy Inference System (ANFIS)

ANFIS combines neural-network learning with fuzzy logic to model nonlinear flood processes and tune fuzzy-inference parameters and structure. It was reported as accurate, generalizable, and easy to implement across hydrological applications.

  • Definition and operation: Fuzzy inference systems approximate human learning with lower complexity and can model nonlinear functions and extreme hydrological events.
  • Definition and operation: ANFIS is an advanced neuro-fuzzy method based on Takagi–Sugeno fuzzy inference that combines ANN learning with fuzzy logic.
  • Definition and operation: ANFIS uses neural learning rules to identify and tune fuzzy-inference-system parameters and structure, including missing fuzzy rules from data.
  • Reported performance: ANFIS became popular in flood modeling because of fast implementation, accurate learning, and strong generalization abilities.
  • Reported performance: ANFIS was used for short-term rainfall forecasts with high accuracy and was reported to offer easier implementation and better generalization with one-pass subtractive clustering.
  • Related decomposition methods: Wavelet transforms decompose time series into resolution bands, improving data quality and reported flood-prediction accuracy and lead times.
  • Related decomposition methods: Wavelet-based neural networks combine wavelet transforms with feed-forward neural networks and were reviewed as capable of highly enhancing model accuracy.

3.5. Support Vector Machine (SVM)

SVMs and SVR provide statistical-learning alternatives to ANNs for flood prediction, using structural risk minimization and feature-space mapping. Reviews report strong generalization and performance, while computational cost and black-box behavior remain concerns.

  • Applications: SVMs predict quantities forward in time by training on past data and can assign new non-probabilistic binary classifiers.
  • Definition and approach: SVMs are supervised learning algorithms based on statistical learning theory and structural risk minimization for flood prediction.
  • Definition and approach: SVMs support linear and nonlinear classification and efficiently map inputs into feature spaces, providing generalization and efficiency.
  • Reported performance: SVMs were reported to provide excellent generalization and better performance than ANNs and MLRs across multiple flood-prediction applications.
  • Reported performance: SVMs are described as more suitable for nonlinear regression problems and for identifying a global optimal solution in flood models.
  • Limitations and variants: SVM limitations include high computational cost, potentially unrealistic outputs, heuristic and semi-black-box behavior, and drawbacks in seasonal-flow prediction.
  • Limitations and variants: LS-SVM improved performance with acceptable computational efficiency by replacing complex quadratic problems with a set of linear tasks.

3.6. Decision Tree (DT)

Decision trees are established flood-prediction tools, including classification and regression variants, while random forests extend them through ensembles of trees. The literature reports strong performance for several tree-based approaches, but applicability remains incompletely investigated in some settings.

  • Decision-tree structure: Decision trees model flood outcomes through branches and target-variable leaves, with classification trees producing discrete class labels.Regression trees instead handle continuous target values, while tree ensembles can reduce final-variable variance.
  • Applications: Decision trees are classified as fast algorithms and became popular in ensemble forms for flood modeling and prediction.The literature also includes CART, REPTs, NBTs, CHAIDs, LMTs, ADTs, and E-CHAIDs.
  • Scope: The applicability of some decision-tree methods to flood prediction has not yet been fully investigated.This scope boundary is stated specifically for DT-based flood modeling.
  • Random forests: Random forests combine multiple decision trees and select among their class predictions as an ensemble.Each tree generates response-predictor values from independent inputs before the ensemble selects classes.
  • Reported performance: Random forests were reported as effective alternatives to SVM, and one comparison found RF delivered the best performance among ANN, SVM, and RF.The M5 decision-tree algorithm was also reported as a strong flood-simulation approach.

3.7. Ensemble Prediction Systems (EPSs)

Ensemble prediction systems combine multiple models to produce automated forecast ensembles and can quantify flood probabilities. The reviewed literature associates EPSs with improved accuracy, robustness, uncertainty handling, and evaluation efficiency.

  • EPS framework: Ensemble prediction systems provide an ensemble of N forecasts, where N is the number of independent realizations of a model probability distribution.EPS models generally use multiple ML algorithms with automated assessment to improve performance.
  • Algorithm composition: EPSs use multiple fast-learning or statistical algorithms, including ANNs, MLPs, decision trees, rotation forests, bootstrap, and boosting.These classifier ensembles allow higher accuracy and robustness.
  • Probabilistic forecasting: Ensemble prediction systems can quantify flood probability from the prediction rate associated with an event.Their quality can be evaluated through verification of the resulting probability distribution.
  • Reported benefits: EPSs were demonstrated to improve model accuracy in flood modeling and were described as efficient prediction systems.The literature also reports timely, automated performance evaluation and management as an EPS advantage.
  • Data processing: Ensemble means, EMD, and EEMD were used with ML methods to improve input quality, dataset management, and flood prediction.EMD-based forecast models nevertheless have documented drawbacks, motivating work on additivity and generalization.

3.8. Classification of ML Methods and Applications

The survey classifies flood-prediction ML methods as single or hybrid and organizes applications by prediction lead time. It finds that popular methods and method characteristics vary substantially across short- and long-term prediction.

  • Method classification: ANNs, SVMs, MLPs, DTs, ANFIS, WNNs, and EPSs are identified as the most popular ML methods for flood applications.These methods are categorized as either single or hybrid approaches.
  • Hybrid methods: Hybrid models combine ML with soft computing, statistical, or physical components whose functions complement individual-method shortcomings.The reported success of these approaches motivated further exploration of advanced hybrid models.
  • Research trend: Figure 4 documents a continuous increase and notable progress in novel hybrid methods over the past decade.This trend supports distinguishing single and hybrid ML prediction models in the survey taxonomy.
  • Prediction horizons: Short-term prediction commonly covers hourly, daily, and weekly horizons, whereas long-term prediction is mainly used for policy analysis.The survey defines a lead time greater than one week as long-term.
  • Survey organization: The characteristics of ML methods vary significantly according to prediction period, making separate short-term and long-term survey categories essential.The study focuses on lead time rather than flood duration or flood type.
  • Review procedure: The survey database records studies by prediction technique, lead-time class, single-versus-hybrid status, and related comparative results.Four tables organize the literature into comprehensive surveys of flood-modeling studies from the past decade.

4. Short-Term Flood Prediction with ML

Short-term flood-prediction studies report strong results for neural, support-vector, tree-based, fuzzy, and hybrid models across different forecasting tasks. Across the review, performance improvements are associated with hybridization, ensembles, optimization, and soft-computing techniques.

  • Neural networks: NARX outperformed BPNN for short-term lead-time prediction and produced an average R2 value of 0.7.The study characterized NARX as effective for urban flood prediction.
  • Single ML methods: MLP was identified as the most accurate model for short-term river-flood prediction.Other comparisons reported ANN, MLP, DT, SVM, SVR, and RF advantages under particular sites, variables, or forecasting conditions.
  • Overall trends: The review reports continuous improvement in speed, complexity, accuracy, and ease of use through ensembles, hybridization, optimization, and soft computing.These strategies are presented as major directions for improving ML flood-prediction models.
  • Hybrid methods: Hybrid models improved prediction accuracy, generalization, uncertainty handling, lead time, or computational performance in multiple reviewed applications.Examples include wavelet, autoregressive, physical-model, optimization, and ensemble integrations.
  • Ensembles and hybrids: Ensemble and hybrid systems were reported as reliable or robust for probabilistic forecasts, flood duration, destructive power, susceptibility, and hourly or daily prediction.Reported examples include EPS, HEC–HMS hybrids, SOM–R-NARX, and WNN-based models.

5. Long-Term Flood Prediction with ML

Long-term flood prediction studies compare diverse single and hybrid ML models across streamflow, rainfall, evapotranspiration, precipitation, and related hydrological tasks. Results favor several specialized models, but performance depends on application and generalization remains a recurring concern.

  • Single-model studies: ANNs, BPNNs, SVMs, MLPs, WNNs, ANFIS, and related models have produced promising long-lead-time forecasts across hydrological applications.The reviewed applications include streamflow, rainfall, evapotranspiration, precipitation, and reservoir inflow prediction.
  • Research scope: Long-term flood prediction remains unresolved because studies have not established which ML method performs best across long lead times.The review summarizes comparative investigations in Tables 4 and 5.
  • ANN applications: ANN models supported long-term evapotranspiration and precipitation forecasting and improved prediction ability over Hilsenhoff’s biotic index using geomorphic data.The ANN application nevertheless exhibited generalization problems.
  • Comparative results: BPNN models were reported as fast and robust for nonlinear flood forecasting, while LLR outperformed BFGSNN on performance and accuracy with larger R2 and lower RMSE.The review also reports BPNN outperforming other methods in one long-term comparison.
  • Comparative results: SVM was a potential candidate for long-term discharge prediction, outperforming ANN in one comparison and being more accurate and easier to build than BPNN and MLR in another.These findings are application-specific comparisons rather than a universal ranking.
  • Hybrid and ensemble models: Hybridization, decomposition, and ensemble approaches repeatedly improved accuracy, stability, generalization, or prediction quality over conventional or single ML models.Examples include WNN-based hybrids, decomposition-enhanced SVR, and hybrid models outperforming ARMA and ARIMA.

6. Comparative Performance Analysis and Discussion

The comparative discussion finds that decomposition, hybridization, ensembles, and optimization generally improve flood-prediction models across accuracy, speed, complexity, generalization, and robustness. However, uncertainty increases with lead time, and models remain data-specific with unresolved generalization and long-horizon limitations.

  • Data decomposition: Wavelet transforms provide multiresolution preprocessing that can improve ML inputs and produce more consistent WNN results than traditional ANNs.The review describes decomposition as a way to expose dataset structure at multiple resolution levels.
  • Data decomposition: Decomposed ML algorithms, including WNN, generally achieved better accuracy, precision, and performance than models trained on undecomposed time series.This pattern was reported for both short-term and long-term rainfall–runoff modeling.
  • Limitations: Longer lead times increased prediction uncertainty, and WNN predictions were not satisfactory at long horizons despite improvements from decomposition and hybridization.The review calls for future evaluation of model precision under increasing lead times.
  • Hybridization: Hybrid decomposition models improved prediction quality, with wavelet–neuro-fuzzy models reported as more accurate and faster than single ANFIS and ANN models.The discussion also links decomposition methods to longer lead times, stability, representativeness, and accuracy.
  • Ensembles: Ensemble prediction systems improved speed, accuracy, and generalization, and some ensemble configurations outperformed traditional ML methods.The reviewed ensembles include bootstrap sampling, genetic programming, Bayesian methods, data fusion, regression, and soft computing techniques.
  • Short-term comparison: Comparative hybrid models showed significant accuracy and generalization improvements over SVM, ANFIS, and ANN for short-term prediction.Figure 10 presents the comparative performance analysis of hybrid ML methods for short-term prediction.
  • Overall conclusions: The review identifies hybridization, ensembles, decomposition, and optimization as the main strategies for improving ML flood prediction, while noting that major models remain data-specific.These strategies are presented as the field’s future direction rather than as a universal model choice.

5. Conclusions

The survey reviews machine-learning models for flood prediction, classifying studies by lead time and by single versus hybrid methods. It identifies four major improvement trends and notes that spatial flood prediction remains outside its scope.

  • The survey analyzed more than 6000 flood-forecasting studies and identified 180 original articles comparing at least two machine-learning models.
  • Models were classified by prediction lead time and then divided into hybrid and single methods for detailed performance comparison.
  • Performance evaluation considered R2, RMSE, generalization ability, robustness, computational cost, and speed across popular methods including ANNs, SVM, SVR, ANFIS, WNN, and DTs.
  • Four major improvement trends were reported: novel hybridization, data decomposition, method ensembles, and optimizer algorithms for model tuning.
  • Ensembles increased model generalization ability and decreased prediction uncertainty, while optimizers helped tune neural-network architectures.
  • Spatial flood prediction was excluded because its modeling methodologies and datasets differ, and it is recommended as a separate future survey topic.

Nomenclatures

This nomenclature section lists abbreviations used for meteorological, hydrological, statistical, machine-learning, and flood-analysis methods.

  • Meteorological and climate abbreviations include WMO, QPF, NWP, GCM, QPE, SPOTA, and POTA.
  • Regression and statistical abbreviations include MNLR, MLR, QRT, FFA, RFFA, LLR, PCA, and SRM.
  • Neural-network and machine-learning abbreviations include ANN, WNN, FFNN, FBNN, MLP, ANFIS, BPNN, SVR, ELM, and SSNN.
  • Time-series and autoregressive abbreviations include ARIMA, ARMA, ARMAX, NARX, WARM, and SAR.
  • Flood-modeling and analysis abbreviations include CHIM, FFRM, HEC–HMS, SFF, MP, SSL, and FFA.
Loading 1908.02781v1…