Paper deep dive
Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment
Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, Robin R. Murphy
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/18/2026, 4:56:56 AM
Summary
This paper presents an empirical investigation into annotator and reviewer performance across multi-source remotely sensed imagery (drone, crewed aviation, and satellite) using the CRASAR-U-DROIDs dataset. Analyzing 74,128 building damage labels from 187 annotators across 9 disasters, the study reveals that revision rates by a consensus committee increase significantly as imagery resolution decreases (from 25.27% for crewed to 36.95% for satellite). It further finds that single-individual reviews reduce but do not eliminate disagreement, leaving higher residual errors in lower-resolution sources. The authors recommend tailoring labeling schemas to sources, prioritizing consensus-based adjudication, and targeting quality control toward lower-resolution imagery.
Entities (17)
Relation Signals (16)
CRASAR-U-DROIDs → contains → Drone Imagery
confidence 98% · The selection of the CRASAR-U-DROIDs dataset [31] was due to its multi-source imagery, comprising drone, crewed, and satellite
CRASAR-U-DROIDs → contains → Crewed Aviation Imagery
confidence 98% · The selection of the CRASAR-U-DROIDs dataset [31] was due to its multi-source imagery, comprising drone, crewed, and satellite
CRASAR-U-DROIDs → contains → Satellite Imagery
confidence 98% · The selection of the CRASAR-U-DROIDs dataset [31] was due to its multi-source imagery, comprising drone, crewed, and satellite
Thomas Manzini → affiliatedwith → Texas A&M University
confidence 95% · Thomas Manzini tmanzini@tamu.edu Texas A&M University
Robin R. Murphy → affiliatedwith → Texas A&M University
confidence 95% · Robin R. Murphy robin.r.murphy@tamu.edu Texas A&M University
Satellite Imagery → hashigherrevisionrate → Consensus Committee
confidence 95% · initial annotations were revised by the final committee at rates that rise steeply from higher- to lower-resolution sources (25.27% for crewed aviation and 36.95% for satellite)
Crewed Aviation Imagery → hashigherrevisionrate → Consensus Committee
confidence 95% · initial annotations were revised by the final committee at rates that rise steeply from higher- to lower-resolution sources (25.27% for crewed aviation and 36.95% for satellite)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed imagery, evaluating human labeling across drone, crewed aviation, and satellite views. Because existing aerial imagery datasets rely predominantly on single-source imagery, there is no currently established state of practice for efficiently allocating human labor to curate large-scale, multi-source aerial datasets. This work addresses this limitation by analyzing annotator and reviewer performance within a post-disaster building damage assessment dataset of 9 disasters, where 20041 buildings in drone, 20695 buildings in crewed aviation, and 33392 buildings in satellite imagery were labeled. These labels, provided by 187 annotators, were then refined through two successive quality-control stages: a single-reviewer pass followed by a consensus-committee review. Our analysis reveals two findings that raise questions for standard crowd-sourcing practices. First, initial annotations were revised by the final committee at rates that rise steeply from higher- to lower-resolution sources (25.27% for crewed aviation and 36.95% for satellite), with the same ordering at every observed workflow stage. Second, a single individual review reduced but did not resolve this disagreement: after review, the committee still revised 6.85% of drone, 14.05% of crewed, and 20.86% of satellite labels. These observations suggest that, in workflows like this one, uniform review allocation leaves the most residual disagreement in lower-resolution imagery. Based on this evidence, and consistent with prior work on adaptive task assignment and budget-aware quality control, this paper offers three recommendations for multi-source dataset curation.
Tags
Links
- Source: https://arxiv.org/abs/2608.14942v1
- Canonical: https://arxiv.org/abs/2608.14942v1
Trouble viewing inline? Open PDF directly →
Full Text
57,735 characters extracted from source content.
Expand or collapse full text
Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment Thomas Manzini tmanzini@tamu.edu Texas A&M University College Station, Texas, USA Priyankari Perali perali@umd.edu University of Maryland College Park, Maryland, USA Raisa Karnik raisak@tamu.edu Texas A&M University College Station, Texas, USA Stephen Johnson sjohnson440@tamu.edu Texas A&M University College Station, Texas, USA Robin R. Murphy robin.r.murphy@tamu.edu Texas A&M University College Station, Texas, USA Synopsis This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed imagery, evaluating human labeling across drone, crewed aviation, and satellite views. Because existing aerial imagery datasets rely predominantly on single-source imagery, there is no currently established state of practice for efficiently allocating human labor to curate large-scale, multi-source aerial datasets. This work addresses this lim- itation by analyzing annotator and reviewer performance within a post-disaster building damage assessment dataset of 9 disasters, where 20,041 buildings in drone, 20,695 buildings in crewed aviation, and 33,392 buildings in satellite imagery were labeled. These labels, provided by 187 annotators, were then refined through two successive quality-control stages: a single-reviewer pass followed by a consensus-committee review. Our analysis reveals two findings that raise ques- tions for standard crowd-sourcing practices. First, initial annotations were revised by the final committee at rates that rise steeply from higher- to lower-resolution sources (25.27% for crewed aviation and 36.95% for satellite), with the same ordering at every observed workflow stage. Second, a single individual review reduced but did not resolve this disagreement: after review, the committee still revised 6.85% of drone, 14.05% of crewed, and 20.86% of satellite labels. These observations suggest that, in workflows like this one, This work is licensed under a Creative Commons Attribution 4.0 Interna- tional License. 2026 ACM Conference on Human-AI Complementarity and Alignment (HCOMP 2026), September 27–30, 2026, Alexandria, VA, USA, 2026. © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2894-5/2026/09 https://doi.org/10.1145/3834580.3838748 uniform review allocation leaves the most residual disagree- ment in lower-resolution imagery. Based on this evidence, and consistent with prior work on adaptive task assignment and budget-aware quality control, this paper offers three recommendations for multi-source dataset curation: (1) tai- lor labeling schemas to each imagery source, (2) prioritize consensus-based adjudication over individual review alone, and (3) preferentially target quality-control effort toward lower-resolution sources. CCS Concepts • Applied computing→Computers in other domains;• Computing methodologies→Computer vision; Ma- chine learning;• General and reference→Empirical stud- ies. Keywords Datasets, Quality Control, Crowd-Sourcing, Annotator Per- formance, Reviewer Performance, Aerial Imagery, Remote Sensing, Human-AI Complementarity, Computer Vision ACM Reference Format: Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, and Robin R. Murphy. 2026. Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment. In 2026 ACM Conference on Human-AI Complementarity and Alignment (HCOMP 2026), September 27–30, 2026, Alexandria, VA, USA. ACM, New York, NY, USA, 13 pages. https://doi.org/10.1145/3834580.3838748 1 Introduction Remote sensing is increasingly used within operational and societal settings [38,42]. Aerial imagery collected from re- mote sensors (satellite, crewed aircraft, and drones) offers wide-area spatial views of scenes that can be used for more arXiv:2608.14942v1 [cs.CV] 14 Aug 2026 EngageCSEdu. https://doi.org/10.1145/3834580.3838748Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, and Robin R. Murphy Figure 1: Example image of tile displayed within the annotation tool, LabelBox. All annotators provided an- notations through LabelBox, where they labeled each building polygon (shown in green and red) within the image tile. informed decision-making. In recent years, human-AI sys- tems have been developed to provide rapid assessments of such imagery for societal and scientific applications, such as disaster response [17,31,47], climate monitoring [13], agriculture [35], and conservation [3]. However, the fun- damental bottleneck in developing these robust models is the human labor required to manually inspect and annotate overhead imagery with prior work relying on crowd-sourced human intelligence spanning drone [5,31,39–41,53], crewed [31, 39, 47], and satellite [4, 14, 17, 25, 29, 31] imagery. Developing robust human-AI systems for remote sens- ing requires large-scale, multi-source datasets [16], yet how varying imagery sources impact the humans tasked with labeling them has been unstudied. Effectively measuring an- notator performance requires data collected from the three major remote sensing aerial imagery sources (drones, crewed aircraft, and satellites) labeled under a unified schema and consistent view type (e.g., nadir). While recent efforts have produced multi-source imagery datasets [8,9,15,39], they fail to meet these criteria: they lack one of the three imagery sources, utilize disjointed schemas across sources, or mix view types (e.g., oblique vs. nadir). Although DeepFlood [9] provides all three sources, the satellite component consists of non-visual data not labeled by annotators. This leaves CRASAR-U-DROIDs [31] as the only known dataset that possesses all three imagery sources independently labeled under a unified schema and view type. As a result, the anno- tation and review data collected during the development of CRASAR-U-DROIDs becomes the necessary subject of this work. Curating such large-scale multi-source datasets to sup- port CV/ML efforts is resource-intensive, where reviewing every annotated label for quality control is often impracti- cal. Along with crowd-sourcing, prior efforts have employed strategies to provide a level of quality control, such as ran- dom sampling the annotations for review [17] or by an ex- haustive review of all annotations, by either one reviewer [40,41] or a committee of reviewers [31]. Consequently, the two-stage review pipeline utilized during the curation of the CRASAR-U-DROIDs dataset provides a rare empiri- cal testbed. Rather than treating this workflow merely as a rigid curation mechanism, it allows for successive quality- control stages, a single-reviewer verification pass followed by a consensus-committee review, to be observed and quan- tified at scale on the same labels. Recent findings highlight that curating multi-source im- agery cannot rely on a direct, 1:1 mapping of labels across sources, as different sources (e.g., drone vs. satellite) alter the perceived content of the imagery [32]. Consequently, each imagery source requires its own dedicated human annota- tion and review pipeline. Despite this massive requirement for independent human labor, there is no prior work inves- tigating how annotator and reviewer performance actually differ across these sources. Without this understanding, there is a limited understanding of the human phenomena that drive the curation of multi-source datasets. This drives the motivating question for this work: How can annotator and re- viewer efforts be targeted for efficient curation of high-quality large-scale multi-source aerial imagery datasets? Guided by this research question, this work presents the first and largest known multi-source analysis of annotator and reviewer performance. Utilizing the CRASAR-U-DROIDs [31] dataset, this work analyzes annotator and reviewer performance across 20,041 drone-derived, 20,695 crewed- derived, and 33,392 satellite-derived building labels observed at three workflow stages: initial annotation, initial review, and committee review. This analysis represents the core con- tribution of this work, yielding two findings: initial annota- tions were revised by the final committee at rates that surge as source resolution coarsens (25.27% for crewed aviation and 36.95% for satellite), and a single individual review left substantial residual revision behind (6.85% for drone, 14.05% for crewed, and 20.86% for satellite). Based on this evidence, this paper suggests three practices for multi-source dataset curation: (1) tailor labeling schemas to each imagery source, (2) prioritize consensus-based adjudication over individual Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage AssessmentEngageCSEdu. https://doi.org/10.1145/3834580.3838748 review alone, and (3) preferentially target review capacity toward lower-resolution sources. 2 Background & Related Work Curating robust datasets for remote sensing requires a deep understanding of human computation, yet annotator and reviewer performance across multi-source imagery remains largely unstudied. To contextualize this gap, this section first outlines the growing reliance on crowd-sourcing to over- come the human labor bottleneck in remote sensing (Section 2.1) and the parallel shift toward multi-source data fusion (Section 2.2). The post-disaster damage assessment is exam- ined as an operational case study to highlight the disjointed quality control strategies currently used in aerial imagery curation (Section 2.3). Finally, prior analyses of annotator performance are reviewed and the necessity of evaluating consensus-based review mechanisms across diverse sensing platforms is established (Section 2.4). 2.1 Crowd-sourcing and Human-AI Systems in Remote Sensing The development of robust human-AI systems in remote sensing is fundamentally constrained by the availability of labeled data. While modern remote sensors collect vast vol- umes of aerial imagery, transforming this raw data into ac- tionable ground truth requires massive amounts of manual human annotation. This requirement creates a severe labor bottleneck, slowing the training and deployment of down- stream computer vision models [52]. To overcome this bottleneck, the remote sensing commu- nity increasingly relies on human computation and crowd- sourcing [11,19,44]. By distributing annotation tasks to large pools of non-expert workers or volunteers, researchers can rapidly scale the labeling of aerial datasets for tasks such as land cover classification [12] and large-scale spatial feature tagging [27]. However, overhead aerial imagery introduces severe per- ceptual challenges not present in standard, ground-level computer vision tasks. Annotators must interpret nadir (top- down) perspectives, arbitrary orientations, variable ground sample distances, and sensor-specific artifacts [52]. These observational complexities make crowd-sourced aerial anno- tations inherently prone to noise and high variance. Conse- quently, deploying crowd-sourced labor for aerial imagery necessitates rigorous quality control and review workflows to ensure the resulting datasets are reliable enough for model training [11]. 2.2 Multi-Source Remote Sensing To build more robust human-AI systems, the remote sensing community is increasingly transitioning toward multi-source data fusion, integrating observations from satellites, crewed aircraft, and drones [16]. By leveraging multiple sensing platforms, downstream models can overcome the temporal and spatial limitations inherent to any single source. This operational shift has driven the recent curation of large- scale multi-source datasets. For example, the FLAIR dataset incorporates multi-source optical imagery for country-scale land cover semantic segmentation [15], while other efforts have generated multi-source datasets for coseismic landslide mapping [8] and inundated vegetation segmentation [9]. While these efforts advance the availability of diverse remote sensing data, they lack the necessary structure to evaluate human annotator performance across differing im- agery sources. Existing multi-source datasets typically fail to meet three strict criteria required for comparative annota- tor analysis: they lack visual data across all three primary platforms, they utilize disjointed labeling schemas across sources, or they mix fundamentally different view types, such as oblique versus nadir perspectives. For instance, while DeepFlood [9] provides multi-source data, its satellite com- ponent consists of non-visual data that human annotators do not manually label. Consequently, there is a critical ab- sence of multi-source datasets featuring independent human annotations collected under a unified schema and consis- tent view type. This structural deficit obscures how differing imagery sources inherently impact the human labor tasked with curating them. 2.3 Quality Control in Aerial Imagery Labeling Post-disaster building damage assessment provides one of the most mature operational testbeds for examining crowd- sourced annotation workflows in remote sensing. Twelve post-disaster building damage assessment datasets contain- ing drone, crewed aviation, or satellite imagery were identi- fied from prior literature [4,5,14,17,25,29,31,39–41,47,53]. Crucially, ten of these are strictly single-source datasets [4,5,14,17,25,29,40,41,47,53]. Only two contain multi- source imagery: Volan v.2018 [39] (drone and crewed) and CRASAR-U-DROIDs [31] (drone, crewed, and satellite). In all known aerial imagery damage assessment datasets, human annotators perform the initial labeling. Annotator populations typically belong to three distinct groups: in- house research teams [25,29,39], recruited domain experts or trained volunteers [4,47,53], and student or crowd-worker populations [31]. Despite this reliance on human labor, there is no stan- dardized consensus on quality control and review strategies. Prior curation efforts typically employ at least one reviewer to inspect labels, but the exact mechanism varies widely. For example, the xBD dataset [17] utilized a two-tiered system EngageCSEdu. https://doi.org/10.1145/3834580.3838748Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, and Robin R. Murphy where all labels received a basic check for obvious errors, followed by an expert review of a 4% random sample. Other datasets mandate an exhaustive review of all annotations by earth observation experts [4]. Several datasets utilize "in- house" reviewers [25,29,31,39–41], often enforcing a single- reviewer pass where annotations are iteratively rejected or approved [40,41]. The LADI v2 dataset [47] instead enforced quality control dynamically through inter-annotator agree- ment, requiring at least three annotators per image. Finally, CRASAR-U-DROIDs [31] utilized a hybrid 2-stage review process consisting of a single-reviewer pass followed by a consensus-committee review. This disjointed landscape highlights a critical gap in the literature: there is no documented analysis of how effec- tively these human reviewers perform, particularly across multi-source aerial imagery. Without understanding the fail- ure rates of single-reviewer versus consensus mechanisms across different sensing platforms, future efforts to efficiently expand and curate operational multi-source datasets will re- main hindered by unverified assumptions about human labor allocation. 2.4 Prior Analyses of Annotator Performance The broader computer vision and natural language commu- nities have increasingly recognized the necessity of rigorous dataset annotation quality management, demonstrating that relying on unverified crowd-sourced labels severely degrades model performance [22,36]. Prior studies emphasize that effi- cient, statistical quality estimation and inter-annotator agree- ment metrics are critical for identifying noise in large-scale datasets [23,36]. However, these generalized frameworks rarely account for the unique spatial and observational com- plexities inherent to remote sensing. Within the remote sensing domain, three relevant prior analyses have investigated annotator performance [2,50,51]. Wang et al.[50] evaluated confidence-weighting mechanisms for crowd-sourced annotations to reduce the need for expert intervention. Similarly, Wang et al.[51] analyzed 50 annota- tors labeling land cover in satellite imagery, finding that ded- icated training improves accuracy and that inter-annotator agreement is a strong indicator of reliability. Most relevantly, Blushtein-Livnon et al.[2] analyzed the performance of 25 annotators on object detection tasks across four annotation strategies. Crucially, their analysis found that a majority- vote annotation strategy outperformed independent expert review efforts. A separate line of quality-control research treats anno- tation as an allocation problem: adaptive task assignment routes items and workers by estimated difficulty and skill [18], selective relabeling asks when an existing label merits another look [28,48], and budget-aware designs optimize labeling and review under a fixed budget [6,21]. Structured adjudication of expert disagreement has been studied directly [46]. However, these mechanisms have rarely been examined in production-scale, multi-source remote sensing workflows, which is the setting this work observes. While the remote sensing analyses above suggest that con- sensus mechanisms can outperform individual reviews, they are limited to single-source imagery. It remains unstudied whether these findings hold across large-scale, multi-source aerial imagery datasets, and how annotator and reviewer performance shifts across sensing platforms. This work ad- dresses this gap empirically. 3 Approach This analysis considers 20,041 drone-, 20,695 crewed-, and 33,392 satellite-derived damage labels provided within the CRASAR-U-DROIDs dataset [31]. Due to its unique feature of containing coincident buildings across three imagery sources (drone, crewed, and satellite), the dataset was selected to be the most appropriate for analyzing revision rates across imagery sources and review stages. This section will detail the data analyzed (Section 3.1), the annotator population (Section 3.2), and the labeling workflow (Section 3.3) used for the dataset’s curation. 3.1 Data The selection of the CRASAR-U-DROIDs dataset [31] was due to its multi-source imagery, comprising drone, crewed, and satellite (an example is shown in Figure 2), and large scale, providing 74,128 building damage labels. This dataset faithfully represents drone-, crewed-, and satellite-based im- agery, as there is a stratification of the image resolutions collected in practice as compared to theoretical resolutions [30, 32]. With the data sourced from CRASAR-U-DROIDs, the an- notation and review data were obtained from a 2-stage label- ing workflow, providing an opportunity to analyze annota- tion and review quality for each source of imagery, which is more extensive than any other prior dataset. Annotator and reviewer statistics and performance were derived from a combination of annotations and metrics stored within Label- Box [24] and the dataset. All labels analyzed were initially provided by annotators via the annotation tool, LabelBox [24]. Only building damage labels that were present in all review stages were considered, and any “manual" and “bulk" buildings were discarded [31]. The building damage assessment labels were provided for post-disaster imagery sourced by either drone, crewed, or satellite, with ground sampling distances at approximately Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage AssessmentEngageCSEdu. https://doi.org/10.1145/3834580.3838748 Figure 2: Example of building across three sources of imagery from the CRASAR-U-DROIDs dataset[31] following Hurricane Ian. [Left] View of building from Maxar satellite imagery (30cm/px). [Middle] View of building from NOAA crewed aerial imagery (15cm/px). [Right] View of building from FL-UAS1 uncrewed (Drone) aerial imagery (3.35cm/px). Notice changes in resolution and appearance across sources. 3cm/px, 15cm/px, and 30cm/px, respectively. The drone im- agery was sourced from the Center for Robot-Assisted Search and Rescue (CRASAR) and collected from nine federally de- clared disasters, comprising six hurricanes (Ian, Idalia, Ida, Michael, Harvey, Laura), the Mayfield Tornado, the Kilauea Volcano Eruption, and the Musset Bayou Fire. The crewed im- agery was sourced from NOAA[37] and was collected from Hurricanes Ian, Ida, Idalia, Laura, Harvey, and Michael. The satellite imagery was sourced from the MAXAR Open data portal [33] and was collected from Hurricanes Ian, Idalia, Harvey, and Michael. Of these, the coincident labels ana- lyzed in this work span eight of the nine events for drone imagery, six for crewed aviation, and three for satellite 1 . 3.2 Annotator Population 187 total annotators participated in the collection of building damage labels and are considered in this analysis. Of the 187 total annotators, 172 were high-school students, 7 were middle-school students, and 8 were undergraduate or gradu- ate students associated directly with the research effort who also served as reviewers. The 179 high-school and middle-school students who vol- untarily participated in the annotation effort did so through STEM outreach events organized to teach students about how machine learning systems are trained and used. Stu- dents from 5 high schools and 1 middle school participated in the effort. Students participating were compensated via educational guest lectures on machine learning and “commu- nity service hours," which were reported to their instructors at the end of the semesters in which students participated. 1 Drone imagery from Hurricane Harvey, and Satellite imagery from Hurri- cane Idalia was annotated in “bulk" by the dataset curation team instead of the annotator pool, was thus excluded from this analysis Among the total 187 unique annotators, 55 labeled drone imagery, 67 labeled crewed aircraft imagery, and 93 labeled satellite imagery. 22 annotators labeled more than one source of imagery, driven by expressed interest by the annotators, or through participation in the research team. It is worth noting that this annotator population differs from the traditional populations found in crowdworking platforms such as Amazon’s Mechanical Turk [1]. The popu- lation of Mechanical Turk Workers is well studied [7,20,43] and differs from the population considered in this work pri- marily in age and education. Prior work has found that 54% of Mechanical Turk workers are between 21-35 years old [26], with 55% of Mechanical Turk workers not having a bachelor’s degree [26]. This represents the primary differ- ence compared to the population considered in this work, of whom 96% have not earned a bachelor’s degree, and under 21 years of age. 3.3 Labeling Workflow The curation of the data in this work employed a 2-stage review process for 43,223 images of 74,128 buildings across 3 imagery sources (drone, crewed, satellite), annotated by 187 annotators for building damage assessment according to the Joint Damage Scale (JDS)[17]. This process is shown in Figure 3. The labeling workflow consisted of five primary components: input data, preprocessing, annotation, post- processing, and review. Differences between this population and the population traditionally used in crowd-working tasks [26,43,45] suggested that additional oversight was needed to ensure annotation quality. The input data consisted of orthomosaics and building polygons. An orthomosaic refers to a large spatial area map, EngageCSEdu. https://doi.org/10.1145/3834580.3838748Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, and Robin R. Murphy Figure 3: Annotation and review workflow used to annotate 43,223 Images of 74,128 building views across 3 imagery sources (Drone, Crewed, Satellite). where each pixel has a longitude and a latitude. The multi- source imagery, detailed in Section 3.1, was provided as or- thomosaics. The building polygons were extracted from the Microsoft Building Footprints dataset [34]. The preprocessing steps consisted of image tiling and over- laying building polygons. Drone orthomosaics were tiled into 2048×2048-sized tiles, and crewed and satellite orthomosaics were tiled into 256×256-sized tiles. The difference kept the ground area covered per tile within a comparable range. The orthomosaics were tiled with a 5% overlap and overlaid with building polygons, resulting in 43,223 image tiles. As the image tiling process was grid-based, buildings were frequently divided across multiple image tiles, as shown in Figure 1. This division of buildings across multiple tiles re- sulted in the 74,128 buildings being presented to annotators in the form of 153,178 unique sub-polygons. These 153,178 sub-polygons would later be recombined for the 74,128 build- ing polygons. There were three annotation opportunities: “Initial An- notation," “Initial Review", and “Final Committee Review." During Initial Annotation, the 153,178 sub-polygons shown in the 43,223 image tiles representing the 74,128 buildings were annotated by 187 annotators through the annotation tool, LabelBox [24]. Initial Annotation began with an ap- proximately 7-minute instructional session detailing how to annotate the building sub-polygons according to the JDS and the mechanics of how to use the LabelBox annotation tool [17]. The same instructional session was provided to the annotators for all three sources of imagery, with only the examples used differing depending on the source of imagery. Immediately after this session, annotators began providing labels. This consisted of a single annotator viewing an image tile and its associated sub-polygons, and providing labels for all the sub-polygons in that image tile. For approximately 45 minutes after this instructional session, the organizers pro- vided real-time feedback on the usage of the LabelBox tool and corrected misconceptions about the application of the JDS. Once this period had elapsed, annotators were permitted to provide annotations without supervision. The Initial Review consisted of a single reviewer inspect- ing the labels provided within LabelBox. These reviewers would either reject, edit, or approve the annotations. If the reviewer rejected a tile, it would be requeued and sent to a different annotator. Reviewers could also choose to manu- ally correct tiles. This review stage was completed once all image tiles were approved. Across the corpus, this stage was carried out by a pool of three to four distinct reviewers per source, including two research-team members in each; each tile received one reviewer. Postprocessing followed in order to prepare the annota- tions for the Final Committee Review. All annotations pro- vided by the 187 annotators were postprocessed to recom- bine and overlay the annotations over the original imagery. Postprocessing was the recombination of the 153,178 sub- polygons into the 74,128 building polygons. This was per- formed by tabulating all of the labels for the sub-polygons and coalescing their labels. In cases where the labels for all sub-polygons agreed, the label was simply that one label. In Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage AssessmentEngageCSEdu. https://doi.org/10.1145/3834580.3838748 Percentile0.50.750.900.950.99 Drone2.03.56.7510.528.0 Crewed1.292.56.1310.3626.5 Satellite1.663.337.0411.9530.92 Figure 4: CDF of seconds annotators spent labeling per polygon across drone, crewed & satellite imagery. Observe that median per-polygon annotation times are similar across imagery sources. Note the x-axis log scale. Percentiles are shown in seconds. cases where there was disagreement, the highest damage la- bel was selected, which biases pre-committee building labels toward higher damage by construction. The Final Committee Review followed and consisted of a committee of reviewers inspecting the reconstructed imagery and labels and correcting any errors. For drone imagery, the fused annotations were overlaid over the original orthomo- saics. For crewed and satellite imagery, the fused annotations were overlaid on the original imagery, but provided within review cards 2 . This stage of review was complete once all annotations were approved by the review committee, thus finalizing the annotations. 4 Analysis The analysis conducted in this work seeks to answer the motivating question introduced in Section 1. This analysis considers both the performance characteristics of the anno- tators within the LabelBox tool and their agreement with the recombined building-level reference labels (Section 5.1). 2 An example Review Card is shown in Appendix B. 4.1 Annotation Timing As discussed, within the LabelBox annotation tool, 187 anno- tators provided labels for 74,128 buildings, which were rep- resented through 153,178 building sub-polygons displayed within 43,223 image tiles. At the median, Initial Annotation took 2 seconds per drone tile, 4 seconds per crewed tile, and 8 seconds per satellite tile. It should be noted that the La- belBox tool records annotation time as an integer number of seconds. However, the differing image resolutions meant that a differing count of buildings was shown in each tile depending on its pixel dimensions. To measure the time spent annotating per building sub- polygon within a given tile, the integer number of seconds spent annotating is divided by the number of building sub- polygons within the tile. When doing this, the trend observed at the tile level disappears. At the median, drone imagery takes 2.0 seconds per building polygon, crewed imagery takes 1.29 seconds, and satellite imagery takes 1.66 seconds 3 . The complete distribution for drone, crewed, and satellite labeling time per building sub-polygon is shown in Figure 4. This finding suggests that there is little difference in the time spent annotating sub-polygons across resolutions. 4.2 Label Revision Rates This work reports revision rates: the fraction of buildings present at both compared stages whose damage label differs. The Final Committee Review labels serve as the workflow’s determinative reference, not independently verified ground truth (Section 5.1). As shown in Figure 5, revision rates rose from drone to crewed to satellite imagery in every stage com- parison, ranging from 6.85% to 36.95%; drone comparisons involving Initial Annotation could not be computed. The error bars in Figure 5 give the 95% Wilson confidence intervals, treating buildings as independent samples, reflect- ing how the labels were elicited, since each building was presented to annotators and reviewers as its own labeling decision. Labels do share context (a single annotator labeled many buildings, different disasters impact different buildings, and satellite buildings recur across views), so these intervals capture sampling variability under a building-level indepen- dence assumption and may understate uncertainty from such shared factors (Section 5.1); annotator effects in particular cannot be modeled separately from source effects, as anno- tator cohorts across sources (Section 3.2) were not consis- tent. Revision rates differed significantly between sources within every stage pair. After individual review, the odds of a committee revision for a satellite label were 3.6×those of a drone label (odds ratio 3.6, 95% CI [3.4, 3.8]) and 1.6×those 3 The statistical significance of the differences between sources cannot be reliably computed due to the integer nature of the timing information recorded by LabelBox. EngageCSEdu. https://doi.org/10.1145/3834580.3838748Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, and Robin R. Murphy Figure 5: Label revision rates between workflow stages (Initial Annotation, Initial Review, and Final Committee Review). Whiskers show 95% confidence intervals under a building-label independence assumption. Rates that could not be computed (pre-review drone labels were not preserved) are marked “unknown". Note the higher revision rates against Initial Annotation and for lower-resolution sources (e.g., Satellite). of a crewed label (95% CI [1.5, 1.7]). Moving beyond confi- dence intervals, it is argued that Stuart-Maxwell marginal- homogeneity tests are the most appropriate in this context because successive stages relabel the same buildings. As a result, distributional shifts between stages are tested with Stuart–Maxwell marginal-homogeneity tests and all found to be푝< .001. Further, statistical significance was also ob- served via chi-squared tests, as all푝< .001 with Cramér’s푉 = 0.11–0.16. When an individual review changed an annotator’s la- bel, the committee endorsed the reviewer’s label in 83.0% (crewed) and 77.2% (satellite) of cases, reverting to the origi- nal annotation in only 8.7% and 10.4%. Most committee revi- sions instead fell on buildings the review had left unchanged (13.5% and 20.3% of which were revised). In this workflow, individual reviewers primarily diverged from the committee by not intervening rather than by intervening differently. 4.3 Label Revision Directionality Figure 6 shows the direction of revisions at each observed stage pair. Four of the seven comparisons were dominated by increases: relative to the labels under review, the Final Committee Review predominantly raised damage severity in all three sources (59.1% of its direction-changing revisions for drone, 75.3% for crewed, and 63.2% for satellite). The Initial Review stage differed by source: near-balanced for crewed (48.2% increases) and predominantly lowering for satellite (74.3% decreases), driven by demotions of minor- and major-damage annotations toward no damage. Overall, the determinative labels sat above the crewed Initial Annotations (62.1% of directional revisions increased severity) but below the satellite Initial Annotations (56.1% decreased): annotators under-reported damage in crewed imagery and over-reported it in satellite imagery, so the direction of annotator bias is itself source-dependent. The mechanisms behind these source-dependent directions can- not be isolated in this observational design (Section 5.1). 5 Discussion The analysis in Section 4 finds that revision rates signifi- cantly differed between imagery sources and review stages, higher for lower-resolution sources, and largest against Ini- tial Annotation, implicating three areas of consideration for curating multi-source aerial imagery datasets. This section first discusses limitations (Section 5.1), the three implications with recommendations for improvements (Section 5.2), and ethical considerations (Section 5.3). 5.1 Limitations There are four limitations of this work, based on the availabil- ity of imagery, damage labels, and their associated annotator and reviewer data. First, this work inherently evaluates varying imagery sources rather than varying sources and ground sample distances. Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage AssessmentEngageCSEdu. https://doi.org/10.1145/3834580.3838748 Figure 6: Direction of label revisions between workflow stages. Under- or over-estimation of damage is measured by the changes made to the building labels at each step of review. Updates to building polygon labels that resulted in a change to or from a label that is not associated with damage (obscured, un-classified) were not included. While evaluating synthetically downsampled or upsampled imagery would isolate the variable of pixel density, it would fail to capture the atmospheric interference, distinct sen- sor artifacts, and off-nadir angle distortions unique to the imagery source. Therefore, the dynamics observed here rep- resent the compounded effects of true multi-source imagery, rather than isolated resolution degradation. Beyond sensing differences, the sources also differ in disaster coverage and the annotator cohorts. As a result, the statistics reported here reflect these factors jointly, and no single cause, including resolution, can be isolated. Second, this work treats the labels arrived at following the final committee review as the determinative label, and reports revision rates against them; they are not indepen- dently verified ground truth. This is notably different from labels generated by an expert inspector at the building site on the ground [10], and prior work has found differences in the distribution of labels generated from different sources [32]. Consequently, a revision measures divergence between workflow stages, not error against an external standard. Third, as mentioned in Section 4, the initial drone labels, prior to their correction by reviewers, are not preserved in LabelBox. This was only recognized after all initial reviews had been completed. The data was later manually preserved for the satellite and crewed aircraft imagery annotations, which were labeled and reviewed afterward. Fourth, the imagery considered in this dataset, and the annotation task itself, is limited to hurricanes, a volcanic eruption, a tornado, and a wildfire. This means that it is unknown if the dynamics observed in this work will persist in other imagery of different disaster types or annotation tasks. 5.2Implications and Recommendations for Annotation and Reviews Across Imagery Sources The findings of the analysis presented in Section 4 imply three items for future data annotation and review efforts for computer vision and machine learning datasets with potential improvements. The first implication concerns allocation: revision rates in this workflow varied widely across sources, so uniformly allocating a fixed amount of review effort left the largest residual disagreement in the lower-resolution sources. This evidence leads to the recommendation that review capacity is better targeted toward the sources that accumulate the most revisions. Second, if label revisions concentrate in lower-resolution sources and go uncorrected, models trained on such data may inherit source-dependent biases; this risk is anticipated here but not measured, and benchmarking it is left to future work. To reduce it, labeling schemas should anticipate each EngageCSEdu. https://doi.org/10.1145/3834580.3838748Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, and Robin R. Murphy imagery source, since the guidance annotators need appears to differ by source. Third, individual review alone did not align labels with the final committee outcome. For crewed and satellite imagery, the committee revised 14.05% and 20.86% of post-review la- bels, respectively, fewer than the 25.27% and 36.95% revisions of the initial annotations, so individual review moved labels toward the eventual committee outcome. However, a substan- tial residual remained, concentrated in buildings the review had left unchanged (Section 4.2). Because the committee outcome is itself the reference (Section 5.1), these residuals measure how far a single review leaves labels from the work- flow’s determinative label, not reviewer error against ground truth. The remaining disagreement after Initial Review under- scores the limits of a single-review stage process. Individual reviewers assign labels without the opportunity to discuss uncertainty when faced with ambiguous imagery, whereas a consensus committee can deliberate on complex edge cases [48,49]; deliberation is also precisely why the committee outcome serves as the workflow’s determinative reference rather than an independent measurement. 5.3 Ethical Implications for Label Collection Crowd-sourcing labels can cause potential negative impli- cations for downstream operations. This work raises three key ethical considerations when crowd-sourcing labels for multi-source datasets: label variation can induce bias and de-calibrate models, constraints on downstream uses, and the need for expert review of annotations to mitigate errors introduced by untrained annotators. Due to a lack of subject-matter expertise in crowd-sourced annotators, labels created by crowd workers may have a large spread of variance. This variance can introduce bias into downstream training and cause de-calibration, restrict- ing potential deployments for trained models and degrading their reliability for operational uses. These concerns moti- vate expert review: consolidating and examining labels lets experts catch misclassifications and dampen crowd-induced variance, and adjudicating rather than authoring labels can limit the bias experts introduce. 6 Conclusion This work contributes the first and largest known empiri- cal analysis of human annotator performance across multi- source remotely sensed imagery. By evaluating 74,128 an- notated buildings labeled by 187 annotators across drone, crewed aviation, and satellite perspectives, this study reveals source-dependent revision patterns within a large-scale an- notation workflow. While the evaluated data originates from post-disaster building damage assessments, the observed dynamics raise data-curation questions relevant to other multi-source remote sensing efforts under the limitations discussed in Section 5.1. The analysis has yielded two surprising findings. First, revision rates rose steeply from higher- to lower-resolution sources: the committee revised 25.27% of initial crewed- aviation annotations and 36.95% of satellite annotations, and the same ordering held at every observed workflow stage. Second, a single individual review reduced but did not resolve this disagreement, leaving 6.85% (drone), 14.05% (crewed), and 20.86% (satellite) of labels to be revised by the consensus committee. These disparities pose risks of model bias and unreliability when such datasets train downstream computer vision models. To mitigate these risks, this evidence suggests that curation efforts move away from uniform review alloca- tion and consider three strategies: (1) tailor labeling schemas to each imagery source, (2) prioritize consensus-based adju- dication over individual review alone, and (3) preferentially target review capacity toward lower-resolution sources. Future efforts will extend this work in three specific di- rections. First, revision rates among buildings annotated by multiple individuals, driven by overlapping image tiles, will be evaluated to measure baseline inter-annotator agreement prior to review intervention. Second, a fine-grained analy- sis of revision rates for specific class labels across sources will be conducted to identify which annotation targets are most susceptible to resolution degradation. Finally, high- performance computer vision models will be benchmarked directly against these human-curated labels to evaluate the operational boundaries of human-AI complementarity in remote sensing applications. Acknowledgments This work is supported by the AI Research Institutes Program funded by the National Science Foundation under the AI Insti- tute for Societal Decision Making (NSF AI-SDM), Award No. 2229881, and under “Datasets for Uncrewed Aerial System (UAS) and Remote Responder Performance from Hurricane Ian” Award No. 2306453. The authors thank the Center for Robot-Assisted Search and Rescue for access to the details of the drone data used in this work. The authors thank the US National Oceanic and Atmospheric Administration (NOAA) for releasing crewed aircraft imagery and Vantor (previously MAXAR) for releasing satellite imagery for these events. Fur- ther, the authors thank the annotators who participated in this work, specifically, the instructors and students at Winch- ester Thurston School, Bryan Collegiate High School, The Galveston Independent School District, Beaver Valley Inter- mediate Unit, Rudder High School, and Bryan Independent School District. Finally, the authors thank Jayesh Tripathi Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage AssessmentEngageCSEdu. https://doi.org/10.1145/3834580.3838748 for his support during the review of the satellite and crewed aircraft data. References [1]Amazon Mechanical Turk. 2025. Amazon Mechanical Turk. https: //w.mturk.com Accessed: 2025-07-17. [2]Roni Blushtein-Livnon, Tal Svoray, and Michael Dorman. 2025. Per- formance of human annotators in object detection and segmentation of remotely sensed data. IEEE Transactions on Geoscience and Remote Sensing (2025). [3]Jeannine Cavender-Bares, Fabian D Schneider, Maria João Santos, Amanda Armstrong, Ana Carnaval, Kyla M Dahlin, Lola Fatoyinbo, George C Hurtt, David Schimel, Philip A Townsend, et al.2022. Inte- grating remote sensing with ecology and evolution to advance biodi- versity conservation. Nature Ecology & Evolution 6, 5 (2022), 506–519. [4]Hongruixuan Chen, Jian Song, Olivier Dietrich, Clifford Broni-Bediako, Weihao Xuan, Junjue Wang, Xinlei Shao, Yimin Wei, Junshi Xia, Cuiling Lan, et al.2025. BRIGHT: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response. Earth System Science Data Discussions 2025 (2025), 1–51. [5] Chih-Shen Cheng, Amir H Behzadan, and Arash Noshadravan. 2021. DoriaNET: A visual dataset from Hurricane Dorian for post-disaster building damage assessment. DesignSafe-CI (2021). [6]Peng Dai, Daniel Weld, et al.2010. Decision-theoretic control of crowd- sourced workflows. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 24. 1168–1174. [7] Djellel Difallah, Elena Filatova, and Panos Ipeirotis. 2018. Demograph- ics and dynamics of mechanical turk workers. In Proceedings of the eleventh ACM international conference on web search and data mining. 135–143. [8]Chengyong Fang, Xuanmei Fan, Xin Wang, Lorenzo Nava, Hao Zhong, Xiujun Dong, Jixiao Qi, and Filippo Catani. 2024. A globally dis- tributed dataset of coseismic landslide mapping via multi-source high- resolution remote sensing images. Earth System Science Data 16, 10 (2024), 4817–4842. [9] Mulham Fawakherji, Jeffrey Blay, Matilda Anokye, Leila Hashemi- Beni, and Jennifer Dorton. 2025. DeepFlood for inundated vegetation high-resolution dataset for accurate flood mapping and segmentation. Scientific Data 12, 1 (2025), 271. [10]Federal Emergency Management Agency. 2025. Preliminary Damage Assessment Guide. PDF report, July 1, 2025. https://w.fema.gov/ sites/default/files/documents/fema_rd_pda-guide_07012025.pdf Effec- tive for incident periods from July 1, 2025; accessed 2025-07-17. [11]Steffen Fritz, Cidália Costa Fonte, and Linda See. 2017. The role of citizen science in earth observation. 357 pages. [12] Steffen Fritz, Ian McCallum, Christian Schill, Christoph Perger, Linda See, Dmitry Schepaschenko, Marijn Van der Velde, Florian Kraxner, and Michael Obersteiner. 2012. Geo-Wiki: An online platform for improving global land cover. Environmental Modelling & Software 31 (2012), 110–123. [13]Yingchun Fu, Zhe Zhu, Liangyun Liu, Wenfeng Zhan, Tao He, Huan- feng Shen, Jun Zhao, Yongxue Liu, Hongsheng Zhang, Zihan Liu, et al. 2024. Remote sensing time series analysis: A review of data and appli- cations. Journal of Remote Sensing 4 (2024), 0285. [14]Aito Fujita, Ken Sakurada, Tomoyuki Imaizumi, Riho Ito, Shuhei Hikosaka, and Ryosuke Nakamura. 2017. Damage detection from aerial images via convolutional neural networks. In 2017 Fifteenth IAPR international conference on machine vision applications (MVA). IEEE, 5–8. [15]Anatol Garioud, Nicolas Gonthier, Loic Landrieu, Apolline De Wit, Marion Valette, Marc Poupée, Sébastien Giordano, et al.2023. FLAIR: a country-scale land cover semantic segmentation dataset from multi- source optical imagery. Advances in Neural Information Processing Systems 36 (2023), 16456–16482. [16]Pedram Ghamisi, Behnood Rasti, Naoto Yokoya, Qunming Wang, Bern- hard Hofle, Lorenzo Bruzzone, Francesca Bovolo, Mingmin Chi, Katha- rina Anders, Richard Gloaguen, et al.2019. Multisource and multi- temporal data fusion in remote sensing: A comprehensive review of the state of the art. IEEE Geoscience and Remote Sensing Magazine 7, 1 (2019), 6–39. [17] Ritwik Gupta, Bryce Goodman, Nirav Patel, Ricky Hosfelt, Sandra Sajeev, Eric Heim, Jigar Doshi, Keane Lucas, Howie Choset, and Matthew Gaston. 2019. Creating xBD: A dataset for assessing building damage from satellite imagery. In Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition workshops. 10–17. [18] Chien-Ju Ho, Shahin Jabbari, and Jennifer Wortman Vaughan. 2013. Adaptive task assignment for crowdsourced classification. In Interna- tional conference on machine learning. PMLR, 534–542. [19]Xiao Huang, Siqin Wang, Di Yang, Tao Hu, Meixu Chen, Mengxi Zhang, Guiming Zhang, Filip Biljecki, Tianjun Lu, Lei Zou, et al.2024. Crowdsourcing geospatial data for earth and human observations: A review. Journal of Remote Sensing 4 (2024), 0105. [20] Panagiotis G Ipeirotis. 2010. Demographics of mechanical turk. [21]David R Karger, Sewoong Oh, and Devavrat Shah. 2014. Budget- optimal task allocation for reliable crowdsourcing systems. Operations Research 62, 1 (2014), 1–24. [22] Jan-Christoph Klie, Richard Eckart de Castilho, and Iryna Gurevych. 2024. Analyzing dataset annotation quality management in the wild. Computational Linguistics 50, 3 (2024), 817–866. [23]Jan-Christoph Klie, Juan Haladjian, Marc Kirchner, and Rahul Nair. 2024. On efficient and statistical quality estimation for data annotation. arXiv preprint arXiv:2405.11919 (2024). [24] Labelbox. 2024. Labelbox. https://labelbox.com [25]Cheng-Chun Lee, Navjot Kaur, Ali Mahdavi-Amiri, and Ali Mostafavi. 2022.Ida-BD: Pre-and post-disaster high-resolution satellite imagery for building damage assessment from Hurricane Ida. Designsafe-CI. Retrieved from [2022-07-25] https://w.designsafe- ci.org/data/browser/public/designsafe.storage.published/PRJ-3563 (2022). [26]Kevin E Levay, Jeremy Freese, and James N Druckman. 2016. The demographic and political composition of Mechanical Turk samples. Sage Open 6, 1 (2016), 2158244016636433. [27]Albert Yu-Min Lin, Andrew Huynh, Gert Lanckriet, and Luke Bar- rington. 2014. Crowdsourcing the unknown: The satellite search for Genghis Khan. PloS one 9, 12 (2014), e114046. [28] Christopher Lin, Daniel Weld, et al.2014. To re (label), or not to re (label). In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing, Vol. 2. 151–158. [29] Szu-Yun Lin, Marisa Edocia, Fang Jung Tsai, Lien an Chen, and Wen-Ni Kuo. 2023. HaitiBRD: A labeled satellite imagery dataset for building and road damage assessment of the 2010 Haiti earthquake. doi:10. 17603/DS2-FQAT-4V02 [30] Jorge-Mario Lozano and Iris Tien. 2023. Data collection tools for post-disaster damage assessment of building and lifeline infrastruc- ture systems. International journal of disaster risk reduction 94 (2023), 103819. [31] Thomas Manzini, Priyankari Perali, Raisa Karnik, and Robin Murphy. 2024. Crasar-u-droids: A large scale benchmark dataset for building alignment and damage assessment in georectified suas imagery. arXiv preprint arXiv:2407.17673 (2024). EngageCSEdu. https://doi.org/10.1145/3834580.3838748Thomas Manzini, Priyankari Perali, Raisa Karnik, Stephen Johnson, and Robin R. Murphy [32]Thomas Manzini, Priyankari Perali, Jayesh Tripathi, and Robin R Mur- phy. 2025. Now you see it, Now you don’t: Damage Label Agreement in Drone & Satellite Post-Disaster Imagery. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. 1998– 2008. [33]MAXAR. 2025. Maxar Open Data. https://registry.opendata.aws/ maxar-open-data/ [34]Microsoft. 2021. Microsoft US Building Footprints. https://github.com/ Microsoft/USBuildingFootprints. [35]Gideon Sadikiel Mmbando. 2025. Harnessing artificial intelligence and remote sensing in climate-smart agriculture: the current strategies needed for enhancing global food security. Cogent Food & Agriculture 11, 1 (2025), 2454354. [36]Joseph Nassar, Viveca Pavon-Harr, Marc Bosch, and Ian McCulloh. 2019. Assessing data quality of annotations with Krippendorff alpha for applications in computer vision. arXiv preprint arXiv:1912.10107 (2019). [37]National Oceanic and Atmospheric Administration. 2025. National Oceanic and Atmospheric Administration (NOAA). https://w.noaa. gov Accessed: 2025-07-17. [38] Ranganath R Navalgund, Vivek Jayaraman, and Parth Sarathi Roy. 2007. Remote sensing applications: An overview. current science (2007), 1747–1766. [39]Yalong Pi, Nipun D Nath, and Amir H Behzadan. 2020. Convolutional neural networks for object detection in aerial imagery for disaster response and recovery. Advanced Engineering Informatics 43 (2020), 101009. [40] Maryam Rahnemoonfar, Tashnim Chowdhury, and Robin Murphy. 2023. RescueNet: A high resolution UAV semantic segmentation dataset for natural disaster damage assessment. Scientific data 10, 1 (2023), 913. [41] Maryam Rahnemoonfar, Tashnim Chowdhury, Argho Sarkar, Debvrat Varshney, Masoud Yari, and Robin Roberson Murphy. 2021. Flood- net: A high resolution aerial imagery dataset for post flood scene understanding. IEEE Access 9 (2021), 89644–89654. [42]Ronald R Rindfuss and Paul C Stern. 1998. Linking remote sensing and social science: The need and the challenges. People and pixels: Linking remote sensing and social science (1998), 1–27. [43]Joel Ross, Lilly Irani, M Six Silberman, Andrew Zaldivar, and Bill Tomlinson. 2010. Who are the crowdworkers? Shifting demographics in Mechanical Turk. In CHI’10 extended abstracts on Human factors in computing systems. ACM SIGCHI, 2863–2872. [44] Ekrem Saralioglu and Oguz Gungor. 2020. Crowdsourcing in remote sensing: A review of applications and future directions. IEEE Geoscience and Remote Sensing Magazine 8, 4 (2020), 89–110. [45]Antonios Saravanos, Stavros Zervoudakis, Dongnanzi Zheng, Neil Stott, Bohdan Hawryluk, and Donatella Delfino. 2021. The hidden cost of using Amazon Mechanical Turk for research. In International conference on human-computer interaction. Springer, 147–164. [46] Mike Schaekermann, Graeme Beaton, Minahz Habib, Andrew Lim, Kate Larson, and Edith Law. 2019. Understanding expert disagreement in medical data analysis through structured adjudication. Proceedings of the ACM on Human-Computer Interaction 3, CSCW (2019), 1–23. [47] Samuel Scheele, Katherine Picchione, and Jeffrey Liu. 2025. LADI v2: Multi-label Dataset and Classifiers for Low-Altitude Disaster Im- agery. In Proceedings of the Computer Vision and Pattern Recognition Conference. 2235–2243. [48] Victor S Sheng, Foster Provost, and Panagiotis G Ipeirotis. 2008. Get another label? improving data quality and data mining using multiple, noisy labelers. In Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining. 614–622. [49] James Surowiecki. 2005. The Wisdom of Crowds. Anchor Books. [50]Xiangyu Wang, Lyuzhou Chen, Taiyu Ban, Derui Lyu, Yifeng Guan, Xingyu Wu, Xiren Zhou, and Huanhuan Chen. 2023. Accurate la- bel refinement from multiannotator of remote sensing data. IEEE Transactions on Geoscience and Remote Sensing 61 (2023), 1–13. [51]Yan Wang, Chenxi Li, Xueyi Liu, Hongdong Li, Zhiying Yao, and Yuanyuan Zhao. 2024. How well do the volunteers label land cover types in manual interpretation of remote sensing imagery? Interna- tional Journal of Digital Earth 17, 1 (2024), 2347443. [52] Gui-Song Xia, Xiang Bai, Jian Ding, Zhen Zhu, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liangpei Zhang. 2018. DOTA: A large-scale dataset for object detection in aerial images. In Proceedings of the IEEE conference on computer vision and pattern recognition. 3974– 3983. [53]Xiaoyu Zhu, Junwei Liang, and Alexander Hauptmann. 2021. Msnet: A multilevel instance segmentation network for natural disaster damage assessment in aerial videos. In Proceedings of the IEEE/CVF winter conference on applications of computer vision. 2023–2032. Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage AssessmentEngageCSEdu. https://doi.org/10.1145/3834580.3838748 Figure 7: An example blank Review Card that was used during the reviews of the Crewed and Satellite imagery. The committee would make corrections to building labels by coloring in the cell corresponding to the up- dated label for the imagery. A Statistics This appendix section details the summary statistics for the building damage annotation task for each of the sources and each of the different data types. This information is shown in Table 1. TilesBuildingsSub-Polygons Drone19,60920,04140,761 Crewed13,12120,69545,582 Satellite10,49333,39266,835 Total43,22374,128153,178 Table 1: A table summarizing the different quantities of items presented to annotators. B Review Cards Review cards were visual representations of the multiple, parallel views of buildings that were used for the Final Com- mittee Reviews for both Crewed and Satellite imagery. These cards captured the pixels from the orthomosaics that rep- resent the building, the building polygon, the labels for those buildings and a grid corresponding to any updates that should be made to the labels for the building labels. An example of these review cards is shown in Figure 7. For crewed and satellite imagery, each review card presented the building’s coincident views from both sources side by side.