Source-linked AI summary
Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment
Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, Robin R. Murphy
TL;DR
Large-scale multi-source aerial datasets require substantial human annotation and review, but prior work provides limited evidence for allocating that labor across imagery sources. This paper analyzes annotator and reviewer performance across drone, crewed aviation, and satellite imagery, finding higher residual disagreement in lower-resolution sources and motivating source-aware curation strategies.
Problem
Existing multi-source datasets lack sufficient independent, comparable human annotations across major imagery platforms, limiting understanding of how to allocate labor for efficient curation.
Method
The paper analyzes 74,128 building labels from drone, crewed aviation, and satellite imagery across initial annotation, single-reviewer verification, and consensus-committee review stages.
Results
Revision rates increased as source resolution decreased, with committees revising 25.27% of initial crewed-aviation annotations and 36.95% of satellite annotations; residual revisions after individual review were 6.85% for drone, 14.05% for crewed, and 20.86% for satellite labels.
Takeaways & Limitations
The findings suggest moving away from uniform review allocation and considering source-specific strategies for multi-source dataset curation.
Takeaways & Limitations
Because imagery sources also differ in atmospheric effects, sensor artifacts, viewing angles, disaster coverage, and annotator cohorts, no single cause, including resolution, can be isolated.
Abstract
from arXiv · showhide
This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed imagery, evaluating human labeling across drone, crewed aviation, and satellite views. Because existing aerial imagery datasets rely predominantly on single-source imagery, there is no currently established state of practice for efficiently allocating human labor to curate large-scale, multi-source aerial datasets. This work addresses this limitation by analyzing annotator and reviewer performance within a post-disaster building damage assessment dataset of 9 disasters, where 20041 buildings in drone, 20695 buildings in crewed aviation, and 33392 buildings in satellite imagery were labeled. These labels, provided by 187 annotators, were then refined through two successive quality-control stages: a single-reviewer pass followed by a consensus-committee review. Our analysis reveals two findings that raise questions for standard crowd-sourcing practices. First, initial annotations were revised by the final committee at rates that rise steeply from higher- to lower-resolution sources (25.27% for crewed aviation and 36.95% for satellite), with the same ordering at every observed workflow stage. Second, a single individual review reduced but did not resolve this disagreement: after review, the committee still revised 6.85% of drone, 14.05% of crewed, and 20.86% of satellite labels. These observations suggest that, in workflows like this one, uniform review allocation leaves the most residual disagreement in lower-resolution imagery. Based on this evidence, and consistent with prior work on adaptive task assignment and budget-aware quality control, this paper offers three recommendations for multi-source dataset curation.
Synopsis
The paper recommends adapting multi-source dataset curation to imagery resolution and using consensus-based quality control. Lower-resolution imagery retains the most residual disagreement after uniform review allocation.
- Uniform review allocation leaves the most residual disagreement in lower-resolution imagery.
- The paper recommends tailoring labeling schemas to each imagery source.
- The paper recommends prioritizing consensus-based adjudication and targeting quality-control effort toward lower-resolution sources.
CCS Concepts
The paper is classified across applied computing, computer vision, machine learning, and empirical studies.
- The work falls under applied computing in other domains.
- The work is classified under computer vision and machine learning.
- The work is categorized as an empirical study.
Keywords
The paper concerns datasets, quality control, crowd-sourcing, annotator and reviewer performance, aerial imagery, remote sensing, human-AI complementarity, and computer vision.
- The paper focuses on datasets and quality control in annotation workflows.
- The paper examines crowd-sourcing, annotator performance, and reviewer performance.
- The paper addresses aerial imagery, remote sensing, human-AI complementarity, and computer vision.
1 Introduction
Remote sensing increasingly depends on human-labeled, multi-source imagery for human-AI systems, but existing datasets provide limited basis for comparing annotation across sources. This work analyzes a two-stage curation workflow to examine source-dependent annotator and reviewer performance and inform efficient labor allocation.
- 1 Introduction: Remote sensing supports rapid assessments for disaster response, climate monitoring, agriculture, and conservation, but manual annotation creates a human-labor bottleneck.
- 1 Introduction: Existing multi-source datasets often lack all three imagery sources, unified labeling schemas, or consistent view types needed for comparative annotator analysis.
- 1 Introduction: The CRASAR-U-DROIDs dataset uniquely provides independently labeled drone, crewed, and satellite imagery under a unified schema and view type.
- 1 Introduction: The curation workflow provides a rare testbed with successive single-reviewer verification and consensus-committee quality-control stages.
- 1 Introduction: The motivating question is how to target annotator and reviewer effort efficiently when curating high-quality, large-scale, multi-source aerial imagery datasets.
- 1 Introduction: The study analyzes 20,041 drone, 20,695 crewed, and 33,392 satellite building labels across initial annotation, initial review, and committee review.
- 1 Introduction: The paper recommends source-specific labeling schemas, consensus-based adjudication, and preferentially targeting review capacity toward lower-resolution sources.
2 Background & Related Work
Remote-sensing datasets increasingly combine imagery sources, but evidence about how source differences affect human annotation and review remains limited. Prior work documents varied quality-control strategies, while cross-source reviewer performance is largely unstudied.
- Large-scale remote-sensing datasets require substantial manual annotation, creating a labor bottleneck for developing robust human-AI systems.
- Crowd-sourced aerial annotation is challenging because annotators must interpret nadir views, arbitrary orientations, variable ground sample distances, and sensor-specific artifacts.
- Multi-source datasets support data fusion across satellites, crewed aircraft, and drones, but existing datasets often lack comparable visual coverage, unified schemas, or consistent view types.
- Post-disaster damage-assessment datasets predominantly use a single imagery source, with only limited multisource coverage documented in prior literature.
- Quality-control practices vary from random sampling and exhaustive expert review to single-reviewer passes and dynamic inter-annotator agreement.
- The literature therefore lacks documented analysis of reviewer effectiveness across multi-source aerial imagery, leaving human-labor allocation assumptions unverified.
- Prior studies examine annotator performance and consensus mechanisms, but their evidence is limited to narrower settings and does not establish cross-platform performance patterns.
3 Approach
The study analyzes a large, three-source post-disaster damage dataset using a two-stage annotation and review workflow. The data include independently labeled drone, crewed-aircraft, and satellite imagery processed into building polygons for quality analysis.
- 3.1 Data: The analysis uses CRASAR-U-DROIDs, a large dataset with drone, crewed, and satellite imagery selected for coincident buildings across sources.The dataset provides 74,128 building damage labels and stratified image resolutions representative of practice.
- 3.1 Data: The imagery represents approximately 3cm/px drone, 15cm/px crewed-aircraft, and 30cm/px satellite ground sampling distances.
- 3.2 Annotator Population: 187 annotators contributed labels, including 172 high-school students, 7 middle-school students, and 8 undergraduate or graduate researchers.
- 3.3 Labeling Workflow: The workflow covered 43,223 images and 74,128 buildings across three sources, using the Joint Damage Scale and two review stages.
- 3.3 Labeling Workflow: Grid-based tiling split buildings into 153,178 sub-polygons, which were later recombined into 74,128 building polygons.
- 3.3 Labeling Workflow: Initial annotation used LabelBox after source-specific examples and shared instruction, with real-time feedback during early labeling.
- 3.3 Labeling Workflow: Initial Review used a single reviewer who could reject, edit, or approve tiles, whereas Final Committee Review corrected reconstructed annotations until approval.
4 Analysis
The analysis compares annotation timing, label revisions, and revision direction across drone, crewed-aircraft, and satellite imagery through successive workflow stages. Revision rates increased for lower-resolution sources, while review-stage changes showed source-dependent directional patterns.
- Annotation Timing: 2 seconds per drone tile, 4 seconds per crewed tile, and 8 seconds per satellite tile were the median annotation times.Tile-level timing differed with imagery resolution and the number of buildings displayed per tile.
- Annotation Timing: 2.0 seconds per drone building polygon, 1.29 seconds per crewed polygon, and 1.66 seconds per satellite polygon were the median normalized annotation times.Dividing tile annotation time by the number of building sub-polygons removed the tile-level trend.
- Label Revision Rates: Revision rates rose from drone to crewed to satellite imagery in every stage comparison, ranging from 6.85% to 36.95%.Drone comparisons involving Initial Annotation could not be computed because pre-review drone labels were not preserved.
- Label Revision Rates: All Stuart–Maxwell marginal-homogeneity tests found stage-wise distributional shifts significant at p < .001.Chi-squared tests were also significant at p < .001, with Cramér’s V = 0.11–0.16.
- Label Revision Rates: After individual review, the committee endorsed changed reviewer labels in 83.0% of crewed and 77.2% of satellite cases.It reverted to the original annotation in only 8.7% and 10.4% of cases, while most committee revisions affected labels the review had left unchanged.
- Label Revision Directionality: The Final Committee Review increased damage severity in 59.1% of drone, 75.3% of crewed, and 63.2% of satellite direction-changing revisions.Initial Review was near-balanced for crewed imagery but predominantly lowered satellite damage estimates, with 74.3% decreases.
- Label Revision Directionality: 62.1% of directional revisions increased severity for crewed Initial Annotations, whereas 56.1% decreased severity for satellite Initial Annotations.The reported direction of annotator bias therefore differed by imagery source, with under-reporting in crewed imagery and over-reporting in satellite imagery.
5 Discussion
The discussion identifies scope and reference-label limitations, then draws recommendations for allocating review and adapting annotation workflows across imagery sources. It emphasizes that individual review reduced but did not eliminate divergence from the committee outcome.
- Limitations: The analysis varies imagery sources together with ground sample distances, so the effects of source and resolution are not separately evaluated.Synthetic resampling would isolate pixel density but would omit source-specific atmospheric, sensor, and viewing-angle effects.
- Limitations: The statistics jointly reflect imagery-source, disaster-coverage, and annotator-cohort differences, so no single cause, including resolution, can be isolated.The observed dynamics represent compounded effects of true multi-source imagery rather than isolated resolution degradation.
- Limitations: Final committee labels are determinative workflow references, not independently verified ground truth, so revision rates measure stage divergence rather than error.This reference-label choice differs from labels generated by expert inspectors at building sites.
- Limitations: Initial drone labels were not preserved before reviewer correction, preventing computation of drone revision rates against Initial Annotation.Satellite and crewed annotations were manually preserved because they were processed later.
- Limitations: The dataset covers hurricanes, a volcanic eruption, a tornado, and a wildfire, leaving persistence of the observed dynamics in other disasters or tasks unknown.The scope boundary applies to both the imagery considered and the annotation task.
- Recommendations: Uniformly fixed review effort left the largest residual disagreement in lower-resolution sources, motivating source-targeted review capacity.The recommendation is to target review toward sources that accumulate the most revisions.
- Recommendations: Uncorrected concentration of revisions in lower-resolution sources may transmit source-dependent biases into trained models, but this risk was not measured.Benchmarking that downstream effect is left to future work.
- Recommendations: Individual review reduced divergence from the committee outcome but left substantial residual disagreement, especially in buildings the review had left unchanged.For crewed and satellite imagery, post-review committee revisions were 14.05% and 20.86%, respectively.
6 Conclusion
This study examines human annotation and review performance across drone, crewed aviation, and satellite imagery, finding substantial source-dependent disagreement. The results motivate source-specific schemas, consensus adjudication, and greater quality-control effort for lower-resolution imagery.
- 6 Conclusion: 74,128 buildings labeled by 187 annotators across drone, crewed aviation, and satellite perspectives support the study’s multisource analysis.The dataset originated from post-disaster building damage assessments.
- 6 Conclusion: The observed dynamics may inform other multi-source remote-sensing curation efforts, although the evaluated data comes from post-disaster building damage assessments.The paper explicitly limits this broader relevance to the limitations discussed in Section 5.1.
- 6 Conclusion: 25.27% of initial crewed-aviation annotations and 36.95% of satellite annotations were revised by the final committee, with the same ordering at every workflow stage.Revision rates rose from higher- to lower-resolution sources.
- 6 Conclusion: 6.85% of drone, 14.05% of crewed, and 20.86% of satellite labels remained subject to committee revision after a single individual review.Individual review reduced but did not resolve source-dependent disagreement.
- 6 Conclusion: The findings suggest that uniform review allocation leaves the most residual disagreement in lower-resolution imagery.The paper connects this observation to adaptive task assignment and budget-aware quality control.
- 6 Conclusion: Future work will assess pre-review inter-annotator agreement, source-specific class-label susceptibility, and human-AI complementarity boundaries.These directions include buildings annotated by multiple individuals and direct benchmarking of computer-vision models against human-curated labels.
A Statistics
The appendix summarizes building-damage annotation quantities for each imagery source and data type. These statistics are presented in Table 1.
- A Statistics: Table 1 summarizes the quantities of items presented to annotators across imagery sources and data types.The appendix identifies the table as the location of these summary statistics.
B Review Cards
Review cards supported final committee reviews of crewed and satellite imagery by combining building views, polygons, labels, and correction choices. For these sources, coincident views were displayed side by side.
- B Review Cards: Review cards displayed orthomosaic pixels, building polygons, labels, and a grid for updating building labels during final committee reviews.The committee made corrections by coloring the cell corresponding to the updated label.
- B Review Cards: For crewed and satellite imagery, each review card presented coincident views from both sources side by side.This side-by-side layout was part of the final committee review workflow.