Source-linked AI summary

FOTBCD: A Large-Scale Building Change Detection Benchmark from French Orthophotos and Topographic Data

Abdelrrahman Moubane

arXiv:2601.22596v1cs.CV

TL;DR

Building change detection lacks large, geographically diverse benchmarks for evaluating geographic domain shift. FOTBCD constructs and releases such a dataset from French orthophotos and topographic data, and experiments show more consistent cross-dataset generalization than geographically constrained training data.

  • Problem

    Existing building change detection datasets provide limited geographic diversity, leaving evaluation under cross-region domain shift insufficiently supported.

  • Method

    FOTBCD aligns temporal BD TOPO differences with BD ORTHO imagery to create binary and instance-level datasets using geographically disjoint French department splits.

  • Results

    Models trained on FOTBCD-Binary generalize more consistently across datasets, while LEVIR-CD+ and WHU-CD models substantially degrade on FOTBCD-Binary.

  • Takeaways & Limitations

    FOTBCD supports robust evaluation under geographic domain shift and future research spanning binary, semantic, and instance-level change detection.

  • Takeaways & Limitations

    FOTBCD is limited to mainland France, so multi-country or cross-continent evaluation requires additional datasets.

Abstract

from arXiv · show

We introduce FOTBCD, a large-scale building change detection dataset derived from authoritative French orthophotos and topographic building data provided by IGN France. Unlike existing benchmarks that are geographically constrained to single cities or limited regions, FOTBCD spans 28 departments across mainland France, with 25 used for training and three geographically disjoint departments held out for evaluation. The dataset covers diverse urban, suburban, and rural environments at 0.2m/pixel resolution. We publicly release FOTBCD-Binary, a dataset comprising approximately 28,000 before/after image pairs with pixel-wise binary building change masks, each associated with patch-level spatial metadata. The dataset is designed for large-scale benchmarking and evaluation under geographic domain shift, with validation and test samples drawn from held-out departments and manually verified to ensure label quality. In addition, we publicly release FOTBCD-Instances, a publicly available instance-level annotated subset comprising several thousand image pairs, which illustrates the complete annotation schema used in the full instance-level version of FOTBCD. Using a fixed reference baseline, we benchmark FOTBCD-Binary against LEVIR-CD+ and WHU-CD, providing strong empirical evidence that geographic diversity at the dataset level is associated with improved cross-domain generalization in building change detection.

1 Introduction

Existing building change detection benchmarks are geographically narrow and limited in scale, motivating FOTBCD’s national-scale, diverse, high-resolution binary benchmark and complementary instance subset.

  • Motivation: Existing benchmarks are constrained by geographic homogeneity and limited scale, reducing coverage of real-world environments and geographic conditions.They may capture region-specific building styles, layouts, vegetation, and imaging conditions rather than broader variability.
  • FOTBCD contributions: FOTBCD-Binary covers 28 departments across mainland France, with 25 for training and 3 held out for evaluation across diverse geographic and urban contexts.The coverage ranges from Mediterranean coastal regions to rural Atlantic areas.
  • FOTBCD contributions: 0.2 m/pixel IGN aerial orthophotos provide high-resolution input imagery for the benchmark.The imagery comes from IGN’s BD ORTHO database.
  • FOTBCD contributions: FOTBCD-Binary supplies pixel-wise CHANGE / NO-CHANGE masks for large-scale building change detection benchmarking.The binary annotations focus on whether building change is present between paired images.
  • FOTBCD contributions: Approximately 28,000 image pairs yield over 100,000 training 256x256 patches, while FOTBCD-Instances adds instance-level polygon annotations.The instance subset distinguishes NEW, DEMOLISHED, and UNCHANGED buildings.
  • Empirical finding: Geographic diversity is associated with stronger cross-domain generalization: FOTBCD-trained models transfer effectively, whereas models trained on LEVIR-CD+ or WHU-CD degrade on FOTBCD-Binary.The comparison uses a fixed reference baseline and includes transfer to WHU-CD and LEVIR-CD+.

2 Related Work

Prior building change detection datasets provide useful benchmarks but remain geographically constrained, whereas FOTBCD emphasizes national-scale diversity and evaluation under geographic domain shift.

  • Existing benchmarks: LEVIR-CD contains 637 image pairs from approximately 20 regions in Texas and is geographically homogeneous despite broad use.Its 1024×1024 Google Earth images have 0.5 m resolution and binary construction-change masks.
  • Existing benchmarks: WHU-CD consists of one high-resolution image pair from Christchurch covering a single earthquake-related change event.The images are 32,507×15,354 pixels at 0.2 m resolution.
  • Existing benchmarks: LEVIR-CD+ adds image pairs to LEVIR-CD while retaining its geographic scope and annotation format.The extension does not remove the underlying regional constraint.
  • FOTBCD: FOTBCD-Binary provides national-level diversity across urban, suburban, and rural environments, while FOTBCD-Instances supplies instance-level polygon annotations.The two releases address complementary benchmark and instance-analysis needs.
  • Geographic generalization: Geographic domain shift causes models to underperform across distinct remote-sensing regions, and dataset-level geographic diversity is proposed as a scalable alternative to adaptation and normalization.Relevant variations include architecture, urban density, land cover, terrain, and imaging conditions.

3 The FOTBCD Dataset

FOTBCD aligns authoritative French orthophotos with topographic building changes, uses department-level geographic holdouts, and releases binary and instance-level annotations for robust evaluation.

  • Data sources: FOTBCD combines IGN BD ORTHO aerial imagery with BD TOPO building-footprint and semantic data.BD ORTHO provides 0.2 m orthophotos, while BD TOPO supplies national topographic building information.
  • Data sources: Temporal differences between BD TOPO snapshots are aligned with BD ORTHO imagery to infer new, demolished, and unchanged buildings.This inference supports both binary masks and instance-level polygons.
  • Geographic coverage: The dataset spans 28 mainland-France departments, using 25 for training and 3 geographically disjoint departments for validation and testing.Department-level splitting prevents geographic overlap between training and evaluation data.
  • Geographic coverage: Selected departments cover dense urban, coastal, rural, mountainous, industrial, and mixed-use environments with varied building typologies and land-cover conditions.The design includes variability in urban layout, surrounding land cover, terrain, and imaging conditions.
  • Annotation schema: FOTBCD-Binary provides pixel-wise binary change masks, while FOTBCD-Instances provides COCO-format polygons with UNCHANGED, DEMOLISHED, and NEW classes.For binary detection, NEW and DEMOLISHED are merged into CHANGE and UNCHANGED buildings are ignored.
  • Annotation schema: Instance polygons support boundary evaluation, instance-level analysis, and future multi-class or instance-aware change detection.Image patches also retain Lambert-93 (EPSG:2154) metadata georeferences.
  • Quality control: Training labels are generated automatically and may contain limited noise, whereas validation and test samples are manually selected and verified.Quality control includes temporal alignment, topological and semantic validation, AI-based filtering, and human verification.

4 HybridSiam-CD: Reference Baseline

HybridSiam-CD is used as a fixed reference baseline for cross-dataset evaluation, combining pretrained semantic features with CNN-based spatial and boundary refinement. Its frozen transformer and standardized training protocol prioritize reproducibility over architectural optimization.

  • Baseline role: HybridSiam-CD combines pretrained Vision Transformer features with CNN-based boundary refinement as a consistent reference model.It is used for benchmarking rather than proposed as a new architecture.
  • Architecture: The siamese design separates semantic processing from spatial processing before fusing both streams into dense change maps.Absolute temporal feature differences encode change, while skip connections and upsampling support decoder fusion.
  • Architecture: The semantic encoder is a frozen DINOv3-sat493M Vision Transformer pretrained on satellite imagery.Freezing the encoder limits trainable parameters and emphasizes stability and reproducibility.
  • Training protocol: All datasets use the same fixed evaluation configuration to ensure comparability across different dataset sizes.The protocol specifies 50,000 optimization steps, cosine annealing with 2,000 warmup steps, combined Lovasz hinge and boundary-aware BCE losses, and standard augmentations.
  • Evaluation and visualization: Table 4 reports IoU with training datasets as rows, evaluation datasets as columns, and diagonal entries representing in-domain performance.The accompanying qualitative visualization grid presents FOTBCD examples as before/after image pairs arranged in rows.

5 Experiments

The experiments assess cross-dataset building change detection under geographic domain shift using IoU. Models trained on geographically constrained datasets transfer poorly to FOTBCD-Binary, whereas FOTBCD-Binary training produces more balanced transfer behavior.

  • Cross-dataset evaluation: The primary experiment evaluates cross-dataset generalization by training on one dataset and measuring IoU on another dataset’s test split.This protocol directly tests performance under geographic domain shift.
  • Generalization to FOTBCD-Binary: 0.3003 and 0.3417 IoU are achieved by LEVIR-CD+ and WHU-CD, respectively, when evaluated on FOTBCD-Binary.These results indicate limited transfer from geographically constrained datasets to the national-scale benchmark.
  • Asymmetric domain shift: Cross-evaluation between LEVIR-CD+ and WHU-CD also produces large drops relative to in-domain training.This supports geographic domain shift as a general challenge rather than a problem specific to FOTBCD-Binary.

6 Discussion

The discussion identifies dataset-level geographic diversity as a key factor associated with stronger cross-region robustness. FOTBCD’s varied French environments improve the breadth of evaluation, while its national scope and building-only labels define important boundaries.

  • Why geographic diversity matters: Geographic diversity at the dataset level is identified as a key factor for improving robustness to cross-region domain shift.The authors relate this observation to broader findings that training-data diversity can matter more than raw dataset size for out-of-distribution performance.
  • Why geographic diversity matters: FOTBCD’s 25 training departments cover multiple climate regimes, architectural styles, settlement patterns, and terrain types.These conditions range from lowland plains to mountainous regions and include oceanic, continental, Mediterranean, mountain, and semi-arid climates.
  • Why geographic diversity matters: Exposure to varied environments encourages models to learn general change-related cues rather than region-specific appearances, layouts, or imaging conditions.The cross-dataset results show constrained benchmarks struggling on FOTBCD-Binary, while FOTBCD training yields more balanced transfer.
  • Limitations: FOTBCD is limited to mainland France, so evaluation across multiple countries or continents requires additional datasets.Its geographic diversity does not establish performance beyond the national scope represented in the benchmark.
  • Limitations: The annotations cover building changes exclusively and omit categories such as roads, vegetation, and land use.Consequently, the benchmark’s labels do not support direct evaluation of those other change categories.

7 Conclusion

FOTBCD is introduced as a national-scale benchmark for building change detection under geographic domain shift, using disjoint French departments and varied environments. Fixed-baseline experiments show more consistent cross-dataset generalization from FOTBCD-Binary than from geographically constrained benchmarks.

  • Conclusion: FOTBCD covers 25 training departments and three geographically disjoint validation and test departments across diverse French environments at 0.2 m spatial resolution.The dataset is derived from authoritative French geographic data and includes urban, suburban, and rural settings.
  • Conclusion: Models trained on geographically diverse national-scale data exhibit improved robustness in cross-dataset evaluation compared with models trained on geographically constrained benchmarks.The conclusion is based on experiments using a fixed reference baseline.
  • Conclusion: FOTBCD-Binary supports large-scale binary benchmarking, while FOTBCD-Instances provides instance-level annotations for broader research directions.The released resources are intended to support semantic change detection and instance-level analysis.

Data Availability and Licensing

FOTBCD separates two publicly available research datasets from a larger nonpublic dataset for commercial use. The public releases provide binary and instance-level building-change annotations, verification splits, and associated research resources.

  • FOTBCD-Binary is publicly released for research use under the CC BY-NC-SA 4.0 license.
  • Approximately 28,000 before/after image pairs include pixel-wise binary building change masks derived from authoritative vector data.
  • Human-verified validation and test splits are included in FOTBCD-Binary.
  • FOTBCD-Instances publicly provides approximately 4,000 before/after pairs with polygons labeled NEW, DEMOLISHED, or UNCHANGED for instance-level evaluation.It complements the binary benchmark and includes dedicated validation and test splits.
  • Both public datasets are distributed through one GitHub repository with training code, configurations, and pretrained checkpoints.
  • FOTBCD-220k contains over 220,000 image pairs but is not publicly released and is distributed separately under a commercial license.Academic access may be granted through collaborative research projects.
Loading 2601.22596v1…