Source-linked AI summary

Combining Physically-Based Modeling and Deep Learning for Fusing GRACE Satellite Data: Can We Learn from Mismatch?

Alexander Y. Sun, Bridget R. Scanlon, Zizhan Zhang, David Walling, Soumendra N. Bhanja, Abhijit Mukherjee, Zhi Zhong

arXiv:1902.01933v1physics.geo-phcs.LGstat.ML

TL;DR

Reliable in situ observations of terrestrial water storage were historically difficult, while GRACE provides an independent information source for evaluating simulated variations. This study combines physically based modeling and deep learning to improve NOAH TWSA, with CNN models improving performance and addressing missing groundwater storage in many parts of the study area.

  • Problem

    Reliable in situ observations made periodic tracking of terrestrial water storage historically difficult, motivating independent information from GRACE.

  • Method

    The study combines physically based modeling and deep learning, using CNN models to predict spatial and temporal TWSA variations and evaluating them with MSE and modified NSE criteria.

  • Results

    CNN models significantly improve NOAH TWSA performance at both the country and spatial scales.

  • Takeaways & Limitations

    The approach addresses missing groundwater storage in NOAH for many parts of the study area.

  • Takeaways & Limitations

    Isolated weak spots remain near the India-Nepal border, particularly on NSE maps, and CNN improvements can be insignificant where NOAH already performs well.

Abstract

from arXiv · show

Global hydrological and land surface models are increasingly used for tracking terrestrial total water storage (TWS) dynamics, but the utility of existing models is hampered by conceptual and/or data uncertainties related to various underrepresented and unrepresented processes, such as groundwater storage. The gravity recovery and climate experiment (GRACE) satellite mission provided a valuable independent data source for tracking TWS at regional and continental scales. Strong interests exist in fusing GRACE data into global hydrological models to improve their predictive performance. Here we develop and apply deep convolutional neural network (CNN) models to learn the spatiotemporal patterns of mismatch between TWS anomalies (TWSA) derived from GRACE and those simulated by NOAH, a widely used land surface model. Once trained, our CNN models can be used to correct the NOAH simulated TWSA without requiring GRACE data, potentially filling the data gap between GRACE and its follow-on mission, GRACE-FO. Our methodology is demonstrated over India, which has experienced significant groundwater depletion in recent decades that is nevertheless not being captured by the NOAH model. Results show that the CNN models significantly improve the match with GRACE TWSA, achieving a country-average correlation coefficient of 0.94 and Nash-Sutcliff efficient of 0.87, or 14\% and 52\% improvement respectively over the original NOAH TWSA. At the local scale, the learned mismatch pattern correlates well with the observed in situ groundwater storage anomaly data for most parts of India, suggesting that deep learning models effectively compensate for the missing groundwater component in NOAH for this study region.

1 Introduction

GRACE provides an independent regional check on model-simulated terrestrial water storage, while physically based and purely data-driven models each have important limitations. The study therefore combines physical modeling with CNNs to learn model–observation mismatch and correct NOAH TWSA over India.

  • GRACE offers an independent source for diagnosing and improving model performance despite its coarse resolution.
  • Model–GRACE disagreement reflects missing storage components, uncertain forcing, limited model capacity, and absent human interventions.
  • Pure data-driven models can lack interpretability, generalizability, and the physical information represented in physically based models.
  • The proposed hybrid approach uses CNNs to learn spatiotemporal mismatch between LSM-simulated and GRACE-observed TWSA, then feeds the learned pattern back to the LSM.
  • After training, CNNs can predict corrected TWSA without GRACE inputs, potentially bridging the GRACE–GRACE-FO gap and reconstructing pre-GRACE TWSA.
  • Over India, all evaluated CNN models significantly improve corrected LSM performance at country and grid scales compared with the original LSM.

2 Data and Data Processing

The study combines Indian groundwater observations with GRACE TWSA and NOAH land-surface-model outputs to evaluate a hybrid correction approach. India provides a relevant test region because groundwater depletion is substantial, while NOAH omits deep groundwater and other storage processes.

  • 2.1 Description of the study area, India: India receives most annual precipitation during the June–September monsoon, averaging 85 cm during that season.
  • 2.1 Description of the study area, India: Groundwater depletion is pronounced in northwest India, where irrigation accounts for over 90% of groundwater use.
  • 2.2 GRACE data: The GRACE mascon TWSA product uses a 0.5°×0.5° grid but represents 3°×3° equal-area caps, with study coverage from April 2002 to December 2016.
  • 2.2 GRACE data: GRACE uncertainty includes measurement and leakage errors, for which the study combines released measurement uncertainty with estimated leakage uncertainty.
  • 2.3 NOAH data: NOAH TWS sums soil moisture, snow water, and canopy water, but omits surface water, runoff routing, deep groundwater storage, and human intervention.

3 Methodology

The methodology represents total water storage as multiple components, defines the NOAH–GRACE mismatch, and trains CNNs to predict that mismatch from gridded predictors. CNN architectures extract multiscale spatial patterns and incorporate predictors, expanded training regions, and transfer learning to address limited geoscience training data.

  • TWS is represented as the sum of snow, canopy, surface-water, soil-moisture, and groundwater storage.
  • The mismatch S(t) is the difference between NOAH-simulated and GRACE-derived TWSA and may reflect systematic or random errors.
  • CNNs learn a functional relationship between mismatch samples and predictor inputs, then predict corrected TWSA without requiring GRACE data.
  • The deep-learning workflow uses GRACE mismatch data for training a regression model rather than directly correcting model states, helping circumvent calibration challenges in conceptually uncertain physical models.
  • CNNs use convolutional layers to extract hierarchical spatial features from gridded inputs, making them suitable for learning multiscale spatial patterns.
  • To improve performance with limited labeled data, the study tests additional precipitation and temperature predictors, broader training regions, and transfer learning.

4 Results

CNN models learned NOAH–GRACE mismatch patterns that improved TWSA agreement across India, with strongest gains in regions where NOAH poorly represents groundwater and precipitation variability. SegnetLite also reproduced GRACE-like country-average dynamics during an unseen period, while some northern and boundary areas remained weaker.

  • Country-level performance: 0.94 average testing correlation exceeded original NOAH–GRACE correlation of 0.83, a 14% improvement.The improvement was statistically significant for all three CNN models, with Williams-test p-values below 0.002.
  • Country-level performance: 0.87 average testing NSE represented a 52% improvement over uncorrected NOAH TWSA.The corrected models better captured observed wet and dry events, although SegnetLite slightly underestimated some 2014 and 2015 wet peaks.
  • Model comparison: SegnetLite achieved the best base-case performance, while Unet and SegnetLite outperformed VGG16 across correlation and NSE distributions.VGG16 likely performed worse because it allowed fewer input feature maps.
  • Spatial patterns: Correction gains were greatest in northwest and south India, where irrigation, groundwater withdrawal, and unresolved precipitation patterns affect NOAH performance.Higher corrected correlation and NSE values were also observed in southcentral and central India.
  • Groundwater comparison: The learned mismatch correlated positively with in situ groundwater storage anomalies across most of India, with a median correlation near 0.4.Correlation was weaker in northwest India, along the India–Nepal border, and in southern coastal areas; surface-water influence weakens the comparison along the Indus and Ganges Rivers.
  • Unseen-period prediction: SegnetLite captured GRACE country-average TWSA during 2017 test months not used in Figure 5a, supporting gap filling between GRACE and GRACE-FO.The testing-period maps showed generally similar improvement patterns, except that correlation correction had little or no effect in north India.

5 Conclusion

The study combines physically based modeling with CNNs to learn and correct NOAH–GRACE TWSA mismatch patterns over India. The approach improves NOAH TWSA at country and grid scales, while indicating broader interpolation and extrapolation potential.

  • The hybrid approach trains VGG16, Unet, and SegnetLite CNNs to learn mismatch patterns between NOAH-simulated and GRACE-observed TWSA, then corrects NOAH TWSA.
  • All considered deep learning models significantly improve NOAH TWSA at both country and grid levels despite a much smaller training sample size.
  • Learned mismatch patterns correlate well with in situ groundwater storage anomalies across many parts of India, suggesting compensation for missing groundwater storage in NOAH.
  • The method offers an alternative for extrapolating TWSA beyond the GRACE period and indicates feasibility for deep-learning-based spatial and temporal interpolation.
  • Relative to conventional data assimilation, the approach is described as using fewer assumptions, extracting multiscale features, and handling multiple data types.
  • The study mainly considers three CNN variants, and its relatively coarse grid resolution leaves finer-resolution testing for future work.

A.1 VGG-16

The VGG16 model uses stacked convolutional and pooling layers to downsample features before producing a 128×128 output through fully connected and linear output layers.

  • VGG16 uses repeated 3×3 convolutional layers with ReLU activation interlaced with 2×2 max-pooling layers.
  • The convolutional features are flattened, passed through a fully connected layer, reshaped to 128×128, and emitted with linear activation.

A.2 Unet model

The Unet model is an encoder–decoder CNN that combines multiscale downsampling and upsampling paths to predict a mismatch field from gridded predictor stacks.

  • Unet is adapted from the original design and belongs to the encoder–decoder class.
  • Its encoder has five downsampling steps, and its decoder has an equal number of upsampling steps.
  • Each downsampling step uses two 3×3 ReLU convolutional layers followed by 2×2 max pooling.
  • Upsampling combines decoder features with same-level encoder maps, and simple upsampling repeats rows and columns without trainable parameters.
  • The model takes image stacks of one or more predictors and outputs the predicted mismatch S(t), using a 1 × 1 convolutional layer with linear activation.

A.3 SegnetLite model

SegnetLite is a reduced Segnet variant using a compact encoder–decoder design, concatenation across matching levels, and trainable transpose convolutions for upsampling.

  • Segnet was introduced for semantic segmentation, while SegnetLite is a smaller variant of the original architecture.
  • Unlike the original design, SegnetLite uses no max pooling and replaces upsampling layers with transpose convolution layers.
  • Transpose convolution layers introduce trainable parameters that learn the upsampling operation.
  • The encoder contains six convolutional layers with filters increasing from 16 to 128, while the decoder is symmetric.
  • Feature maps from encoding and decoding steps are concatenated at the same level before output upsampling and 1 × 1 convolution.

Figure Captions

The figure captions describe CNN inputs, outputs, training use of the observed mismatch, and comparisons among corrected TWSA, NOAH, and GRACE across time and space. They also document seasonal mismatch maps, model architectures, prediction intervals, and correlations with in situ groundwater storage anomalies.

  • Model inputs: NOAH TWSA is the base predictor, with precipitation and temperature as auxiliary inputs stacked across multiple time steps.The input stack includes data at t, t −1, …, t −n.
  • Model operation: The CNN output is predicted S(t) with the same dimensions as the input, while observed mismatch is used only during training.After training, GRACE data are no longer required for prediction.
  • Mismatch patterns: Seasonal maps show the learned mismatch S(t) for DJF, MAM, JJA, and SON, with fitted probability-density functions and colors scaled between (-25, 25) cm.The seasonal panels summarize mismatch distributions and spatial patterns.
  • Model evaluation: Country-level and grid-level figures compare NOAH and CNN-corrected TWSA with GRACE using time series, CDFs, correlation coefficients, and NSE.Training and testing periods are separated in the country-level time series, and GRACE uncertainty is shown with an error bound or 95% prediction intervals.
  • Spatial comparisons: SegnetLite-corrected TWSA is compared with GRACE through maps of TWSA, their differences, and corresponding correlation and NSE maps.For plotting, the correlation and NSE maps are scaled to [-1,1].
  • Groundwater comparison: The learned SegnetLite mismatch is evaluated against in situ groundwater storage anomalies using a correlation map and its coefficient CDF.The map coordinates are grid-cell indices from 0 to 127.
Loading 1902.01933v1…