Source-linked AI summary

A Controlled Evaluation of Model Rankings and Input Reliance in Surface Water Segmentation

Kittipat Phunjanna, Kristóf Karacs, Chayut Ngamkhanong

arXiv:2608.30895v1cs.CV

TL;DR

Surface-water segmentation rankings do not by themselves reveal component contributions, ranking stability, input reliance, or deployment scope. The study evaluates these distinctions through repeated configuration comparisons, paired test-chip analyses, fixed-checkpoint input stresses, geographic reweighting, and a secondary GEOID-Flood evaluation. The cross-modal student leads Sen1Floods11 on three-seed mean IoU, but the evidence shows that ranking, attribution, reliance, and scope require separate analyses.

  • Problem

    Aggregate IoU rankings compare complete configurations but do not establish why systems differ, whether close orderings are stable, or how strongly predictions rely on individual inputs.

  • Method

    The study combines repeated configuration comparisons, paired test-chip analysis, fixed-checkpoint input stress tests, geographic reweighting, and supervised ancillary-input evaluation on GEOID-Flood.

  • Results

    The cross-modal student has the highest Sen1Floods11 three-seed mean, while close rankings vary across seeds and composition; supervised ancillary-input effects mostly agree on GEOID-Flood.

  • Takeaways & Limitations

    Global IoU supports selection among tested configurations, whereas stability, component attribution, input reliance, and deployment claims require distinct evidence.

  • Takeaways & Limitations

    The reported scores concern retrospective all-water segmentation with temporal look-ahead, and configuration comparisons do not isolate individual component contributions.

Abstract

from arXiv · show

Performance evaluation for surface-water segmentation commonly uses an aggregate metric such as global intersection-over-union (IoU) to rank model configurations. However, a configuration ranking does not by itself establish why one system performs better, whether a close ordering is stable, or how strongly predictions rely on individual inputs. We examine these distinctions primarily on Sen1Floods11 through repeated configuration comparisons, paired test-chip analysis, fixed-checkpoint input stress tests, and geographic reweighting, with a targeted secondary evaluation of supervised input configurations on GEOID-Flood. The cross-modal student achieves the highest three-seed mean IoU on Sen1Floods11, but close orderings vary across seeds and geographic weighting, while ancillary-input rankings differ between Swin-UNet and U-Net. The GEOID-Flood evaluation shows substantial agreement in supervised ancillary-input effects, although the exact architecture ordering remains configuration dependent. Fixed-checkpoint tests further establish reliance on terrain and WorldCover without establishing a clean-input performance benefit, while target semantics and the later WorldCover prior restrict the evaluation to retrospective all-water segmentation. These results show that aggregate metrics remain useful for ranking complete configurations, but ranking stability, component attribution, input reliance, and deployment scope require distinct evidence. Performance evaluation should therefore match the evidence reported to the claim being made.

1. Introduction

Surface-water segmentation pipelines combine multiple data, supervision, and optimization choices, so aggregate rankings do not explain individual contributions or ranking stability. The evaluation therefore matches pooling, paired resampling, fixed-checkpoint stress tests, and geographic reweighting to distinct claims.

  • Motivation: SAR, terrain, hydrology, land cover, weak labels, and optical training information jointly determine benchmark scores.A score reflects the complete pipeline rather than an explanation of each input or choice.
  • Motivation: Sen1Floods11 combines permanent and non-permanent water in flood-event chips, while prior gains bundle architecture, labels, supervision, and ancillary inputs.These complete pipelines establish benchmark performance but do not isolate architecture, priors, privileged supervision, or interactions.
  • Evidence distinctions: End-to-end comparisons rank complete systems but cannot identify which changed component caused a difference when initialization, parameterization, and optimization also vary.Training with an ancillary input likewise does not establish that predictions depend on it.
  • Evidence distinctions: Paired test-chip resampling diagnoses composition sensitivity, fixed-checkpoint stress tests establish reliance under an input change, and geographic reweighting tests aggregation dependence.Pooling remains appropriate for ranking complete configurations, but none of these analyses alone identifies component causality or deployment performance.
  • Research questions: The evaluation asks whether close global-IoU rankings remain stable, which attribution and reliance claims are supported, and how semantics, provenance, and geography constrain scope.Its central principle is to match evidence to the claim being made.

2. Related Work

Related work uses SAR and ancillary information for flood mapping, while benchmark reevaluations motivate evidence beyond aggregate scores. Prior studies also show that training-only optical supervision and cross-dataset transfer involve distinct data and optimization regimes.

  • SAR and ancillary information: Sen1Floods11 provides Sentinel-1 imagery with manual and automatically generated surface-water labels for geographically distributed evaluation.Common baselines include U-Net and encoder–decoder descendants, while DEM, HAND, and WorldCover provide terrain, drainage, and land-cover information.
  • Training-only optical supervision: Cross-modal distillation trains a deployable Sentinel-1 student using a Sentinel-1/Sentinel-2 teacher, with Sentinel-2 used only during training.The narrower fixed-teacher block holds teacher outputs constant while varying deployable inputs.
  • Evidence beyond aggregate scores: Benchmark reevaluations separate distribution shift, image-cue reliance, and pipeline comparability rather than treating aggregate rankings as explanations.Scores can vary with initialization, sampling, and pipeline choices, motivating significance and replicability checks.
  • Evidence beyond aggregate scores: High intra-dataset flood-mapping scores need not transfer across datasets, while dataset-level IoU weights regions through their foreground unions.These observations motivate examining both generalization and aggregation behavior.

3. Evaluation Design

The evaluation compares complete architectures and ancillary-input configurations, then probes ranking sensitivity, geographic composition, and fixed-checkpoint input reliance. Its metrics and resampling procedures are designed to distinguish configuration performance from component attribution and broader deployment claims.

  • Benchmark and data: The HandLabeled split contains 252 training, 89 validation, and 90 test chips at 512×512 pixels and 10 m ground sampling.The ten named geographic groups occur across all splits, while a separate 15-chip Bolivia set is excluded.
  • Metrics: Primary foreground IoU pools true positives, false positives, and false negatives over valid test pixels before computing the metric.This pooled calculation differs from an equal-chip mean; F1, precision, and recall are secondary metrics.
  • Architecture comparisons: Six complete architecture configurations are compared under the same five-channel recipe, but initialization and parameter counts are not identical across configurations.Only vanilla U-Net uses random initialization; the other five use ImageNet initialization, so the comparison does not isolate architecture alone.
  • Ancillary-input comparisons: Swin-UNet and vanilla U-Net each use eight VV+VH ancillary-input subsets spanning combinations of DEM, HAND, and Water across seeds 42, 1337, and 2026.Changing a subset also changes the input stem, so these are complete input-configuration comparisons rather than single-channel ablations.
  • Cross-modal pipeline: The cross-modal student uses five non-optical inputs, while fixed-teacher comparisons share teacher outputs but retain differences in student input width and training trajectory.The teacher receives SAR, terrain, Water, and Sentinel-2 bands; the student is trained with frozen soft targets and uses only declared deployable inputs.
  • Input reliance: Input stress tests alter fixed checkpoints by replacing, translating, or perturbing WorldCover, DEM, and HAND while holding other specified conditions fixed.These tests measure sensitivity to selected out-of-distribution changes, not clean-input benefits or naturally occurring failure rates.
  • Paired comparisons: Paired differences preserve model pairing across shared chip resamples, while the three-seed mean summarizes only the evaluated fixed runs.The reported bootstrap ranges describe sensitivity to observed test composition rather than retraining variability, and their interval procedures did not support categorical separation, equality, or equivalence labels.
  • Geographic sensitivity: Geographic reweighting gives equal weight to ten named group-level IoUs, complementing global IoU’s greater influence for regions with larger foreground unions.Region-cluster resampling and leave-one-region-out recomputation diagnose dependence on the released composition, not new-country or new-event performance.

4. Results

The cross-modal student leads the archived Sen1Floods11 ranking, but close orderings vary across seeds and geographic weighting. Configuration comparisons establish rankings of complete systems, while fixed-checkpoint tests separately show terrain and WorldCover reliance without establishing clean-input benefits.

  • Configuration ranking: 0.7045 mean IoU makes the cross-modal student the highest-ranked configuration, while Swin-UNet leads supervised systems at 0.6964.The best individual run is the cross-modal student’s seed-42 IoU of 0.7114; the training-free S1OtsuLabelHand reference reaches 0.5458 IoU.
  • Ranking stability: The cross-modal–Swin-UNet ordering changes by (+1.87), (-0.01), and (+0.59) percentage points across seeds, with paired ranges spanning zero each time.The mean ranking favors the cross-modal student on archived runs, but retraining does not establish the same ordering.
  • Ranking stability: Geographic weighting reverses several rankings: pooled IoU ranks Swin-UNet first, whereas equal-region IoU ranks full U-Net first.Cross-modal–Swin-UNet ranges span zero across seeds, so the average lead describes the released benchmark composition rather than a universal ordering.
  • Component attribution: The six five-channel architecture means occupy a 1.70-point band, and Swin-UNet’s seed differences versus vanilla U-Net have paired ranges spanning zero.Swin-UNet is therefore the highest-ranked member of the controlled group, not evidence of an isolated transformer advantage.
  • Component attribution: Ancillary-input rankings differ by architecture: full input leads Swin-UNet by 1.80 points over VV+VH, while VV+VH+Water leads U-Net by 1.12 points.VV+VH+HAND is consistent for Swin-UNet but not U-Net, supporting architecture-dependent configurations rather than a universal HAND effect.

5. Discussion, Implications, and Limitations

The evaluation separates configuration ranking from evidence about stability, input reliance, and scope. It finds that fixed-checkpoint stresses establish dependence on terrain and WorldCover, while target semantics, provenance, uncertainty, and geography constrain interpretation.

  • Interpretation of rankings: A high score supports selecting a tested configuration, not attributing performance to one architectural or input component.Different channel rankings across Swin-UNet and U-Net prevent a universal channel prescription.
  • Configuration performance and input reliance: Fixed-checkpoint stress tests establish dependence on terrain and WorldCover, but not clean-input benefits or naturally occurring failure rates.Clean-input comparisons identify useful configurations under tested recipes, whereas stresses change inputs while holding checkpoints fixed.
  • Target semantics and data provenance: The reported scores describe retrospective all-water segmentation with temporal look-ahead, not transient-inundation or deployment performance.The target combines permanent and event-specific water, and WorldCover postdates the evaluated SAR acquisitions.
  • Uncertainty and population scope: Three seeds and paired test-chip resampling reveal sensitivity to observed training and test composition but do not characterize the full retraining-outcome distribution.The reported ranges remain descriptive because candidate interval procedures did not pass every empirical coverage check.
  • Implications for benchmark evaluation: Global IoU ranks configurations, repeated and paired analyses assess stability, fixed-checkpoint stresses assess reliance, and semantics, provenance, and geography define scope.Future evaluation should separate transient and permanent water, use temporally valid priors, match training choices in component comparisons, and test more seeds on held-out events.

6. Conclusion

The paper argues that benchmark rankings should be interpreted according to the construct measured and the evidence supporting each claim. Across Sen1Floods11 and a secondary GEOID-Flood evaluation, aggregate scores aid configuration selection, but stability, attribution, reliance, and deployment require separate evidence.

  • Conclusion: Global IoU measures overlap with the Sen1Floods11 all-water target, while GEOID-Flood repeats supervised ancillary-input comparisons.On GEOID-Flood, most supervised ancillary-input changes agree in direction, but cross-modal, reliance, and deployment claims are not evaluated.
  • Conclusion: The cross-modal student has the highest Sen1Floods11 three-seed mean, but ranking alone does not establish stability, component contributions, or input reliance.A higher score aids model selection but neither explains why a system performs better nor establishes performance in another setting.
  • Conclusion: Aggregate metrics rank systems; paired analyses test stability; interventions test reliance; and semantics, provenance, and geography define scope.The central implication is to match the evidence reported to the claim being made.
Loading 2608.30895v1…