Paper deep dive
CHIRP dataset: towards long-term, individual-level, behavioral monitoring of bird populations in the wild
Alex Hoi Hang Chan, Neha Singhal, Onur Kocahan, Andrea Meltzer, Saverio Lubrano, Miyako H. Warrington, Michel Griesser, Fumihiro Kano, Hemal Naik
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/27/2026, 1:37:46 AM
Summary
The CHIRP dataset is a comprehensive, long-term behavioral monitoring resource for wild Siberian jays, designed to bridge computer vision research with biological applications. It supports multiple tasks including individual re-identification (re-id), action recognition, 2D keypoint estimation, object detection, and instance segmentation. The paper also introduces CORVID, a novel pipeline for individual bird identification based on the segmentation and classification of colored leg rings, and proposes application-specific benchmarking metrics like feeding rates and co-occurrence rates to evaluate model performance in real-world biological contexts.
Entities (5)
Relation Signals (3)
CHIRP → containsdatafor → Siberian jay
confidence 100% · The CHIRP dataset is curated from a long-term population of wild Siberian jays
CORVID → performs → Individual Re-identification
confidence 100% · CORVID (COlouR-based Video re-ID), a novel pipeline for individual identification of birds
CHIRP → supportstask → Action Recognition
confidence 95% · supporting re-identification (re-id), action recognition, 2D keypoint estimation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Long-term behavioral monitoring of individual animals is crucial for studying behavioral changes that occur over different time scales, especially for conservation and evolutionary biology. Computer vision methods have proven to benefit biodiversity monitoring, but automated behavior monitoring in wild populations remains challenging. This stems from the lack of datasets that cover a range of computer vision tasks necessary to extract biologically meaningful measurements of individual animals. Here, we introduce such a dataset (CHIRP) with a new method (CORVID) for individual re-identification of wild birds. The CHIRP (Combining beHaviour, Individual Re-identification and Postures) dataset is curated from a long-term population of wild Siberian jays studied in Swedish Lapland, supporting re-identification (re-id), action recognition, 2D keypoint estimation, object detection, and instance segmentation. In addition to traditional task-specific benchmarking, we introduce application-specific benchmarking with biologically relevant metrics (feeding rates, co-occurrence rates) to evaluate the performance of models in real-world use cases. Finally, we present CORVID (COlouR-based Video re-ID), a novel pipeline for individual identification of birds based on the segmentation and classification of colored leg rings, a widespread approach for visual identification of individual birds. CORVID offers a probability-based id tracking method by matching the detected combination of color rings with a database. We use application-specific benchmarking to show that CORVID outperforms state-of-the-art re-id methods. We hope this work offers the community a blueprint for curating real-world datasets from ethically approved biological studies to bridge the gap between computer vision research and biological applications.
Tags
Links
- Source: https://arxiv.org/abs/2603.25524v1
- Canonical: https://arxiv.org/abs/2603.25524v1
Trouble viewing inline? Open PDF directly →
Full Text
57,241 characters extracted from source content.
Expand or collapse full text
CHIRP dataset: towards long-term, individual-level, behavioral monitoring of bird populations in the wild Alex Hoi Hang Chan 123 , Neha Singhal 4 , Onur Kocahan 4 , Andrea Meltzer 1235 , Saverio Lubrano 1235 Miyako H. Warrington 56 , Michael Griesser 12357∗ , Fumihiro Kano 123∗ , Hemal Naik 138∗ 1 Centre for the Advanced Study of Collective Behaviour, University of Konstanz 2 Dept. of Collective Behavior, Max Planck Institute of Animal Behavior 3 Dept. of Biology, University of Konstanz, 4 Dept. of Computer and Information Science, University of Konstanz 5 Luondua Boreal Field Station 6 School of Biological and Medical Sciences, Oxford Brookes University 7 Dept. of Zoology, Stockholm University 8 Dept. of Ecology of Animal Societies, Max Planck Institute of Animal Behavior ∗ shared senior authorship. hoi-hang.chan, andrea.meltzer, saverio.lubrano, michael.griesser, fumihiro.kano@uni-konstanz.de onurkocahan, neha.singhal.blr@gmail.com, mwarrington@brookes.ac.uk, hnaik@ab.mpg.de Abstract Long-term behavioral monitoring of individual animals is crucial for studying behavioral changes that occurs over different time scales, especially for conservation and evo- lutionary biology. Computer vision methods have proven to benefit biodiversity monitoring, but automated behavior monitoring in wild populations remains challenging. This stems from the lack of datasets that cover a range of com- puter vision tasks necessary to extract biologically mean- ingful measurements of individual animals. Here, we intro- duce such a dataset (CHIRP) with a new method (CORVID) for individual re-identification of wild birds. The CHIRP (Combining beHaviour, Individual Re-identification and Postures) dataset is curated from a long-term population of wild Siberian jays studied in Swedish Lapland, support- ing re-identification (re-id), action recognition, 2D keypoint estimation, object detection, and instance segmentation. In addition to traditional task-specific benchmarking, we introduce application-specific benchmarking with biologi- cally relevant metrics (feeding rates, co-occurrence rates) to evaluate the performance of models in real-world use cases. Finally, we present CORVID (COlouR-based Video re-ID), a novel pipeline for individual identification of birds based on the segmentation and classification of colored leg rings, a widespread approach for visual identification of in- dividual birds. CORVID offers a probability-based id track- ing method by matching the detected combination of color rings with a database. We use application-specific bench- marking to show that CORVID outperforms state of the art re-id methods. We hope this work offers the community a blueprint for curating real-world datasets from ethically ap- proved biological studies to bridge the gap between com- puter vision research and biological applications. 1 Introduction Behavior is often the first response of animals when adapt- ing to environmental changes [53, 55]. Measuring changes in behavior over time is hence crucial for conducting re- search in behavioral ecology and conservation. In recent years, rapid developments in technologies have opened new avenues for measuring behavior of wild animals, one of which is the application of computer vision to replace man- ual observation from images and videos. Recent developments in computer vision have shown im- pressive applications for collecting behavior data in field conditions (outdoor), from 2D [16, 24, 32, 60, 63] and 3D posture estimation [12, 30, 57] to action recognition [3, 9, 22, 56] and individual re-identification [20, 33, 52]. However, large-scale or long-term deployment of these ap- plications is still limited mainly due to two problems. Firstly, computer vision research focuses on solving spe- cific tasks such as object detection, re-id or keypoint estima- tion, whereas long-term monitoring often requires multiple computer vision tasks to be solved simultaneously. In long- term monitoring, biologists are generally interested in iden- 1 arXiv:2603.25524v1 [cs.CV] 26 Mar 2026 Figure 1. CHIRP dataset summary. A) Solving the problem of who, with video re-identification datset of individuals. B) Solving the problem of what, with action recognition dataset and 2D keypoint estimation. C) Additional annotations to support the main tasks, including segmentation (yellow) of color rings, bounding box (green box) and segmentation (yellow) of birds. D) Application specific benchmark, 12 independent test videos with per frame annotation for bounding box, identities and behaviors, with novel metrics on errors related to biological measures like individual feeding rates and paired co-occurrence rates. tifying who did what behavior. To achieve automated mon- itoring, re-identification and behavioral recognition tasks have to be solved together, either directly or using sup- porting sub-tasks like detection, tracking and keypoint es- timation. Therefore, the classic strategy of developing task- specific methods is not suitable for behavioral monitoring because often it is difficult to readily combine and deploy novel algorithms directly for data collection. This challenge can be alleviated by curating large datasets with annotations to solve multiple tasks using the same real-world data. Secondly, traditional benchmarking processes lack addi- tional metrics to support deployment for biological appli- cations. While some recently published datasets provide annotations for solving a larger range of computer vision tasks (See supplementary; [17, 35, 36, 45–47]), benchmark- ing is often done for each individual task in isolation. This creates ambiguity for users while selecting methods for de- ployment, because task specific benchmarking does not re- flect the impact of choosing a specific method on the fi- nal biological measurements. The unpredictable nature of error propagation is demonstrated with recent case studies [8, 48], and a possible solution is to include a benchmarking mechanism that allows tasks to be combined and tested. In this paper, we address these problems and introduce computer vision solutions for long-term, individual-level behavioral monitoring in the wild.Firstly, we present the CHIRP dataset (Combining beHaviour Individual Re- identification and Postures), the first dataset focusing on social behaviors in a wild bird population, the Siberian jay (Perisoreus infaustus), a group living corvid. Secondly, we present the application-specific benchmark, a novel evalu- ation paradigm that makes use of application-specific met- rics like feeding rate and co-occurrence rates of individu- als. Our benchmarking allows novel methods to be directly tested for their impact on the final biological measurement, encouraging method optimizations for downstream appli- cations instead of individual tasks. Thirdly, we introduce CORVID (COlouR-based Video re-IDentification), a novel automatic re-id framework as a baseline for the applica- tion specific benchmark, by identifying birds via individual color rings. While color leg rings are widely used to mark individuals in wild bird populations [2, 7], to the best of our knowledge, we present the first computer vision frame- work to automatically identify individual birds using color rings. We hope this work can illustrate how datasets and evaluation procedures can be designed with clear deploy- ment objectives, to allow novel computer vision methods to be better applied for downstream biological understanding. 2 Related Works Over the last decade, the popularity of animal-specific datasets has increased significantly.Initial datasets in- cluded task-specific data to tackle computer vision prob- lems such as detection and tracking [58, 65], 2D and 3D keypoint estimation [16, 30, 34, 38, 64], action recogni- tion [4, 5, 10, 23, 31] and re-identification (see [66]). These datasets have contributed to bringing computer vision inno- vations to animal studies, allow biodiversity monitoring to be scaled and automated [54]. However, it is widely ac- knowledged that long-term studies involve solving multiple computer vision tasks together. In response, recent datasets are moving towards more holistic and real-world datasets that contain multiple computer vision tasks within the same study system (see supplementary). Datasets such as Animal Kingdom [47] primarily focus on the problem of animal behavior recognition of 850 an- imal species, with ground truth labels on action recogni- tion tasks, but also for object detection and keypoint esti- mation. However, the dataset was collated from internet- sourced youtube videos and direct applicability of methods developed with the dataset for biological studies remain un- validated. Datasets like LoTE-animal [35] or Baboonland [17] use data directly sourced from biological studies and focuses on action recognition tasks, with support of a wide range of computer vision tasks but do not provide re-id. Simi- 2 Figure 2. Summary of leg color ring definitions and naming convention of the Siberian jays. A) Sample image of an individual, and a list of the different color rings . Ring positions are defined as top/bottom left/ right from the bird’s perspective. B) Class distribution of ring masks provided in the ring segmentation dataset. C) Naming convention, each bird is named in the order of top left, bottom left, top right, bottom right ring. The color combination of bird in the picture is oaor: orange (o), aluminium (a), orange (o), red (r). larly, datasets like 3D-POP [45] and Bucktales [46], focus on tracking large groups for collective behavior and provide ground truth for bounding box, re-id and 2D/3D postures but no annotations for individual activity. These datasets allow novel algorithms to be better bridged to downstream biological applications because they are curated from the data collected for biological studies. In the CHIRP dataset, we follow the same blueprint to ensure that both behavior and identity annotations are supported for downstream ap- plications. The WILD dataset [62] is a dataset from a social be- havior study and offers multi-view 3D tracking and re-id of cowbirds in semi-wild environment. While not focus- ing on action recognition, that dataset proposes bird re- identification, with the bird subjects also fitted with color leg rings. However, the authors did not explicitly make use of color ring information for re-id, instead use an image classifier on an image crop of the birds. Similarly, Individ- ualBirdID [20] also use cropped image of the back patterns of birds for re-id, showing that CNNs can reliably identify individuals across 3 bird species. Here, we include bird re- id problem from an additional perspective of using color information from the leg rings with CORVID. Finally, ChimpACT [36] is a dataset on semi-captive chimpanzees in Leipzig Zoo, with ground-truth labels on identities, bounding boxes, postures and behaviors, focus- ing on longtitudinal (long-term) monitoring of chimpanzee groups. ChimpACT provides task specific benchmarking results [36, 37], which makes it hard to evaluate whether individual model performances are sufficient to achieve ac- curate longtitudinal individual-level behavioral monitoring. CHIRP dataset includes application specific benchmarks, with new metrics designed to evaluate the performance of algorithms directly in the context of the final use case. 3 The CHIRP Dataset To facilitate long-term, individual-level, behavioral moni- toring, it is critical to record who does what (Figure 1). Thus, we first provide a large-scale video re-identification dataset to determine who is present in the video frame. Then, to classify what the individuals are doing, we provide annotations for action recognition of behavioral classes, and 2D keypoints for fine-scaled kinematics. Below, we first describe the study system and data collection, then detailed descriptions of the provided annotations, and fi- nally the application-specific benchmark that complements the dataset. All data samples in the dataset will be pro- vided with a date of data acquisition, to allow for time- aware splits. We also note that except for the video re-id dataset, all datasets adopt an 80/20 train/test split. The com- plete dataset, code and metadata is available here: https: //github.com/alexhang212/CHIRP_Dataset 3.1 Data collection CHIRP dataset is collected in a long-term study population of Siberian jays (Perisoreus infaustus), in Swedish Lapland (65°40’ N, 19°0’ E) between 2014 and 2022. Siberian jays are group-living birds (group size range: 2-7); groups typ- ically consist of a breeding pair and non-breeder, which are either retained offspring or unrelated immigrants [18]. Groups are year-round territorial and can be found in pre- dictable locations, allowing reliable monitoring of groups. Each bird is fitted with an aluminum ring and a unique com- bination of 2 to 3 colored plastic rings of 11 unique col- ors (permit via Ume ̊ a Animal Ethics Board, A23-20 and A26-13 under the license of the Swedish Museum of Nat- ural History). This provides up to 1331 (11 3 ) possible unique combinations, allowing researchers to visually iden- tify individuals in the field (Figure 2A). Color rings en- 3 sure monitoring accuracy across lifetime and is not affected by changing plumage of birds. Due to the unique social structure of Siberian jays, the long-term study strives to an- swer questions related to the evolution of cooperative be- havior and social buffering effects towards environmental changes [28, 59]. In this study, videos are taken as part of a standardized behavioral recording protocol, where 15-30 min long videos (25fps, 1920 x 1080) are recorded in each group using a standardized feeding perch. Then, researchers manually an- notate identity and behavior from videos using the BORIS annotation software [21] e.g. feeding bouts, submissive be- haviors, and displacements for each individual (see supple- mentary for complete ethogram). BORIS [21] is a widely used annotation software in animal behavior studies to man- ually code time of behavior bouts, and these annotations were used to reduce annotation effort. Overall, 443 unique videos (across 9 years, 2014-2022) were used to prepare the CHIRP dataset, and date for each data sample is always pro- vided for time-aware splits. We also ensured that samples from the same behavioral videos are never assigned in the same split for all subsequent datasets. In addition, we also collected multi-view data on jays foraging on the ground, which only contributes to the 2D keypoint dataset. 3.2 Problem 1: Who is present? 3.2.1 Video re-identification dataset The video re-identification (re-id) dataset consists of 16,190 video clips (25 frames, 1 second each) for 183 unique in- dividuals (Figure 1). Each individual has on average 89 samples, and 32.42% of birds (N= 59) has samples across multiple years. We first train a YOLOv8 object detection model [29] (see subsection 3.4) to detect birds within each 15-30 minute long behavioral video. Then using manual annotation from BORIS, 1-second long cropped video clips (25 frames) are extracted, in video segments where only 1 bird was present for both BORIS and YOLO, allowing auto- matic matching of identities. We manually check the video clips to ensure the automatic ID assignment is correct, filter- ing out 8.3% false positive video clips due to manual mis- labelling and false positive detections in YOLO. We ensure all short clips from the same 15-30 minute behavioral video are always in separate data splits, to assure no data leakage from background information is present. Using the clips of individuals, we prepared a closed split, disjointed split and open split based on definitions formal- ized by Wildlife Datasets [66]. The closed set split is clos- est to the final use-case for the Siberian jay system, as re- searchers are always present when recording new videos al- lowing new individuals to be acknowledged and added to the gallery. In the closed-set split, all individuals are both in the train and test set (with 80/20 ratio), and the task is to assign individuals in the test set to individuals in the train- ing set. Next, the disjointed split is created to test for gen- eralization across systems. We assign approximately half the individuals to the train and test set respectively. Within the test set, we further assign all clips from a single be- havioral video as the gallery, and the rest as query. The re-identification task here is to match query clips to an indi- vidual within the gallery database. Finally, the open-set split is prepared to test for the ro- bustness against new individuals, for alternative deploy- ments that may not involve researchers on site (e.g. pas- sive camera traps). We assign 20% and 80% of individ- uals as ”unknown” and ”known” respectively. Within the known individuals, data samples are split into train and test set (80/20 ratio), with data samples of the unknown indi- viduals added to the test set. The re-identification task is to determine if a data sample is ”known” or ”unknown”, then if known, assign the correct ID. Siberian jays are territorial and group members remain fairly consistent within each location, with occasional en- counters with neighbouring territories [27]. Taking advan- tage of this biological feature, we offer two levels of meta- data for each video. First, a short-list of individuals within the territory (N=2-4, mean 2.92), and second, a list of indi- viduals that are also likely likely to be seen in that territory (neighbours; N=9-25, mean 14.58). This adds a layer of complexity and varying difficulty to the dataset, allowing domain-specific knowledge to be built into potential com- puter vision solutions for the re-id problem. 3.3 Problem 2: What are they doing? 3.3.1 Action recognition dataset The action recognition dataset includes manual annotations of 1387 short video crops of birds (min 25 frames, max 74 frames, mean 66.73 frames) performing each behavior (Fig- ure 1). We code “eat” as a bird that was pecking at the food, “submissive” as individuals displaying stereotyped wing- flapping behavior [27], and “others” as any other behav- ior, which includes vigilance, resting and flying (see Fig- ure 3 for class distribution). To speed up annotation, we first train a YOLOv8 model [29] (see subsection 3.4) to extract bird tracks from each video, then use BORIS annotations to identify and extract video segments that contains behaviors of interest. Finally, we manually review and annotate each short clip with the appropriate behavior. 3.3.2 2D Keypoint estimation dataset We provide a comprehensive 2D keypoint dataset of 1176 individual bird instances, across 879 images (Figure 1), with manually annotated keypoint ground truth for 13 unique keypoints. 36% of frames are taken from standard- ized feeding videos, while 64% are taken from videos from another feeding context, where birds are foraging on the ground. Each instance also includes corresponding bound- 4 Figure 3. Class distribution for action recognition dataset ing box annotation for each individual bird. 3.4 Additional annotations In addition to the main tasks described above, we also provide additional annotations to support the main tasks. Firstly, we provide annotations for object detection and in- stance segmentation. For this, we provide 1156 frames of 1669 bird instances with annotated bounding box and seg- mentation mask of birds and the feeding stick, as well as bounding box around the food. To reduce manual anno- tation time, we use SAM2 [49] to automatically generate masks of birds based on bounding box prompts, which we validate against 688 annotations, with a mean 0.84 inter- section over union (IOU). Secondly, as the color rings are the most visually salient features that directly encodes indi- vidual identities within an image in Siberian jays, we also provide bounding box and segmentation masks of individ- ual color rings. We provide annotations across 944 frames of cropped images of birds, with 2713 unique ring instances of 12 unique color classes (Figure 2B). Finally, to allow for training end-to-end or multi-modal methods, we also provide model-annotated datasets by us- ing current best model (see Section 5) to produce 2D key- points and segmentation for the video re-id and action recognition datasets. 3.5 Application specific benchmark We provide 12 independent test videos of 35 bird individ- uals with frame-by-frame annotations of eating behavior, identities and bounding boxes for the application specific benchmark. These videos are not present in any other an- notations provided, thus no data leakage is possible. We first use YOLOv8 trained on the bounding box annotations above to obtain 2D tracks, then manually assigned bird ID to each track. We match these tracks to additional BORIS annotations, where an observer marks every instance where the beak of a bird individual touches the food when feed- ing. Since the tracks obtained from YOLO contains seg- mented tracks (e.g., when an individual jumps from one side to the other), we use linear interpolation to join two tracks that were marked as the same individual. The re- sulting dataset consists of frame-by-frame bounding box with coded IDs and all instances of feeding by each individ- ual. In addition to acting as an application specific bench- mark, this dataset also acts as a multi-object tracking (MOT) benchmark dataset due to the availability of frame-by-frame bounding box annotations with identities. Since we design the application specific benchmark specifically for the CHIRP dataset, we also introduce novel metrics to evaluate and compare the performance of mod- els for application relevant use cases (see supplementary methods for detailed descriptions). First, we evaluate on two lower-level metrics, 1) proportion correct frame assign- ments, defined as the proportion of ground truth tracks and frames that are assigned to the correct individual, to eval- uate tracking and individual identification performance. 2) We compute mean precision, recall and F1-score for indi- vidual feeding events by splitting each video into 1s time windows, with true positives defined as pecking of the given individual detected within the same time window, averaged across all individuals. Next, we also calculate two higher- level biological measures, 1) individual level feeding rates (pecks/minute) and 2) co-occurrence rates, as the proportion of time spent together between each pair of individuals, di- vided by total video length. For both biological measures, we computed the mean, median and standard deviation of absolute errors and Pearson’s correlation between the pre- dicted and ground truth. 4 CORVID: COlouR-based Video re-ID To set an initial benchmark for the re-identification task, we propose a novel pipeline that aims to detect individual color rings for individual identification. Attaching unique combi- nations of color rings to wild birds is common practice for population monitoring (861 ”combination of uncoded color leg rings” projects covered by cr-birds database [19]) and long-term demographic studies (105/175 populations color ringed covered by SPI-bird database [14]), thus a method for individual identification based on color rings is widely applicable. However, to the best of our knowledge, we are the first to directly leverage this feature of bird study sys- tems to achieve individual identification through the detec- tion of color ring patterns. Previous work uses computer vision methods to automatically detect tags that are placed on animals, including color barcodes [40, 43, 44, 51], and fiducial markers (e.g QR codes or aruco tags; [1, 13, 61]. Other work has also explored deep learning based classi- fiers to recognize bird individuals, both in captivity and in the wild [20, 62]. However, compared to existing re-id ap- proaches, our approach do not rely on any of the training data provided in the video re-id dataset, and only relies on the detection of color rings. This allows for the method to be the generalized to any new individuals given the ring combinations are known, and no new color is introduced. The pipeline has three main components (Figure 4). 5 Input clips Mask2Former instance segmentation Identify ring pairs (distance threshold) a: 0.01 b:0.03 c: 0.00 g: 0.00 l: 0.00 m: 0.01 o: 0.01 p:0.78 r:0.16 s:0.00 w:0.00 y:0.00 a: 0.01 b:0.03 c: 0.00 g: 0.00 l: 0.00 m: 0.01 o: 0.01 p:0.78 r:0.16 s:0.00 w:0.00 y:0.00 a: 0.01 b:0.03 c: 0.00 g: 0.00 l: 0.00 m: 0.01 o: 0.01 p:0.78 r:0.16 s:0.00 w:0.00 y:0.00 a: 0.01 b:0.03 c: 0.00 g: 0.00 l: 0.00 m: 0.01 o: 0.01 p:0.78 r:0.16 s:0.00 w:0.00 y:0.00 Random forest Color identification Resize and convert to hsv space mapg: 0.62 marm: 0.47 Possible birds list Select most likely individual Ring pair color probabilities pooled across frames Figure 4. CORVID pipeline. Schematic for the color based re-ID approach pipeline. 1 second clips from CHIRP are fed into Mask2Former instance segmentation model, to extract masks of rings. The rings are grouped into ring pairs based on a distance threshold, then resized and converted into hsv space. The images are fed into a random forest model to predict probabilities of each color, then combined with associated ring pair to create a probability matrix of every color pair, then pooled across frames. Finally, the most probable bird is selected based on the possible birds that could be present in a given video. Firstly, we detect individual rings using a Mask2Former in- stance segmentation model [11] trained on the ring segmen- tation dataset, then cropped, transformed into HSV space, and resized into a 20x20 resolution images. In the second step, we feed color histograms from the images into a ran- dom forest model trained on the ring segmentation dataset, to output confidence scores for each color. We formalize the problem as a multi-class classification problem to allow for the model to predict confidence scores for each color class, considering some color classes are similar to each other. As the final step, we implement a matching algorithm by first identifying ring pairs based on centroid distance threshold of each ring detection, then sum up the probability of ev- ery paired ring color combination based on the outputs of the random forest classifier, across the 25 video frames. We then match the final score with the possible ID metadata for the data sample, and the most likely ID for a given video clip is identified. We refer to the supplementary section for more details on the pipeline and exploration of the CHIRP ring segmentation dataset. 5 Benchmarking To explore the performance of state-of-the-art models on the CHIRP dataset, we provide task-specific benchmarks for each of the main tasks proposed. Next, we implement a sim- ple pipeline to be applied on the application-specific bench- mark, to provide a baseline on how the best algorithms per- forms when combined for automated data extraction. 5.1 Video re-identification For video re-id, we compare our proposed CORVID pipeline to Mega Descriptor, a foundation model for an- imal re-identification [66]. We compare CORVID with Mega-descriptor pre-trained on other animal datasets, and Mega Descriptor that is fine-tuned with CHIRP (Table 1). For Mega Descriptor, we pooled frame-wise probabilities within a tracklet, and the ID with the highest average score was assigned. For benchmarking, we only benchmarked the closed and disjointed split, as we do not have a reliable way of distinguishing between known and unknown individu- als using our proposed CORVID method. For each data split, we test three conditions, we use possible birds for the given data sample in the gallery (“within territory”), possi- ble birds with neighbours (”within territory + neighbours) and all birds in the re-id dataset (”All”). We found that CORVID outperforms MegaDescriptor with the within ter- ritory and neighbours constraint, but the contrary is true in the closed-set when the gallery include all individuals (Ta- ble 1). This shows that explicitly detecting and using ring color information from the CORVID pipeline seem to be better than state-of-the-art deep metric learning techniques, but this relies on the within-territory constraint. 5.2 Action Recognition Next, we train a series of models for video action recog- nition using the MMAction2 library [41]. We benchmark video-based methods here, but posture based methods can also be used. We train each model for 100 epochs, and the best epoch in terms of top 1 accuracy in the test set was chosen (Table 2). Overall, C3D performs the best, with an accuracy of 0.72. 5.3 2D Keypoint Estimation For keypoint estimation, we train models using the M- Pose library [42]. We compute mean and median Euclidean error, root mean squared error (RMSE) and percentage cor- rect keypoints (PCK05, PCK10), as a keypoint estimate that lies within 5% and 10% of the largest dimension of the ground truth bounding box. We train all models for 100 epochs, and the best epoch based on test PCK is selected. Table 3 shows benchmarking results, with ViTPose large being the best performing model, and all architectures per- forming well as evident from the high PCK values. This is indicative that the keypoint annotations provided will al- low for training of accurate keypoint estimation models for further tasks like action recognition. 6 Table 1. Video Re-ID benchmarks. We compare CORVID, pre-trained and fine-tuned mega descriptor across different evaluation settings. Bold denotes best performing model for each metric. Method Closed setDisjointed set Within Territory Within Terr. + Neighbours All Within Territory Within Terr. + Neighbours All Top-1Top-3Top-1Top-3Top-1Top-3Top-1Top-3Top-1Top-3Top-1Top-3 CORVID0.660.960.290.490.050.070.690.970.310.530.060.13 Pre-trained Mega0.280.670.190.410.100.190.310.620.140.320.050.10 Fine-tuned Mega0.270.560.170.350.100.170.410.710.130.270.050.09 Table 2. Action recognition benchmarks. We compute precision, recall, f1-score and top-1 accuracy for each model. Bold denotes best performing model for each metric. ModelPrecisionRecallF1Accuracy SlowFast0.4600.6780.5480.678 C3D0.6750.7150.6840.715 X3D0.4600.6780.5480.678 Table 3. 2D Keypoint estimation benchmarks. Mean, median and root mean squared error (RMSE) was computed using Euclidean distances be- tween predicted and ground truth keypoints. PCK10 and PCK05 stands for percentage correct keypoints within 10% and 5% of ground truth, scaled by the largest dimension of the bounding box. Bold denotes best performing model for each metric. ModelMean Error (px)Median Error (px)RMSE (px)PCK@10PCK@5 ResNet-5012.037.56820.650.9400.832 ResNet-10112.477.49121.240.9400.830 ResNet-15210.917.24517.940.9610.848 HRNet10.866.46319.420.9510.862 ViTPose-small10.326.87216.700.9570.859 ViTPose-large7.7735.19412.640.9780.915 5.4 Application-specific benchmark We implement a simple pipeline that combines different models and algorithms with the aim of extracting individ- ual co-occurrence and feeding rates. The pipeline includes 1) object detection to identify bounding boxes of birds, 2) tracking algorithm to combine detections into tracklets, 3) individual identification for each tracklet and 4) action recognition within each tracklet. As an initial experiment to determine how accuracies in task-specific metrics propa- gates to application-specific metrics, we compared 3 meth- ods for ID assignment: CORVID, MegaDescriptor fine- tuned on the disjointed set and random assignment (re- peated 100 times, to obtain mean estimates). For all three methods, we use a list of possible birds within the terri- tory to constraint the possible birds to a list of 2-5 individ- uals. As solutions improve in the future, the possible bird list can be expanded to include neighbours, and eventually the whole population. For the rest of the pipeline, we use YOLOv8 for object detection and BoTSORT (using boxmot library; [6]) for tracking. For action recognition, we use the best performing C3D model for action recognition by split- ting tracklets into 1s segments. In addition, we also provide a human benchmark, by coding the first 5-minutes of each video by an independent human coder to provide baseline value, as the target for future models to reach. We find that differences in performance between CORVID and Mega Descriptor in task-specific metrics (Ta- ble 1) predicts the accuracy in application-specific metrics, as evident from higher performance when using CORVID in the pipeline (Table 4, 5). Surprisingly, random assign- ment performs the best in some metrics (Table 5), showing that the proposed methods still have room for improvement. We also find that MegaDescriptor performs worse than ran- dom assignment across all higher-level metrics (Figure 5), highlighting that the model is not suitable for deployment. Finally, errors from all pipelines are still high compared with the human benchmark in both individual feeding rates and co-occurrence rates (Figure 5, Table 5), and further im- provements will be needed, highlighting the value of the CHIRP dataset. Table 4. Lower-level baseline metrics for application specific bench- mark. Metrics evaluate tracking, ID assignment accuracy, and behavioral recognition performance. Bold denotes best performing model for each metric. Individual Identification Prop. correct frames Peck Precision Peck Recall Peck F1 CORVID0.6470.4850.7250.537 MegaDescriptor0.6170.4100.5500.408 Random0.3310.3030.4360.327 Table 5. Higher level biologically relevant metrics. Comparison of indi- vidual feeding rates and co-occurrence rates across different pipeline con- figurations. Bold denotes best performing method for each metric, and arrows represents whether a higher or lower value is better Individual ID Individual feeding ratesCo-occurrence rates Mean↓Median↓SD↓r↑Mean↓Median↓SD↓r↑ CORVID9.0008.1317.3420.5820.0410.0190.0570.654 MegaDesc.13.148.60413.080.5050.0560.0280.0630.557 Random9.3499.2875.6670.4370.0410.0280.0460.799 Human1.880.803.470.9100.0280.0070.0530.913 6 Discussion and Limitations Application driven ML is increasingly important, which raises recent discussions on how computer vision innova- tions can be bridged to real world applications [8, 50]. The CHIRP dataset is a task-diverse dataset consists of exist- ing data collected from an ongoing long-term system. This ensures that the dataset is directly relevant for automated 7 Figure 5. Application specific benchmark results. Comparing ground truth measurements and predictions from pipelines to test for how different components affects biological measurements. We compared proposed CORVID pipeline, fine-tuned MegaDescriptor and random assignments for individual recognition. All pipelines used YOLOv8 for object detection, BoTSORT for tracking, and C3D for action recognition. Absolute errors of A) individual-level feeding rates and, B) co-occurrence rates and correlation of C) individual-level feeding rates and, D) co-occurrence rates. Individual feeding rates defined as number of times individual pecks at the food (pecks/min), and co-occurence rates is defined by the proportion of time two individuals were detected together, scaled by video length. individual-level behavioral monitoring in the wild. To fur- ther the goal of bridging computer vision and the appli- cation domain, we also introduce the application-specific benchmarks, a collection of independent test videos, with novel metrics to evaluate the ability of computer vision al- gorithms to extract biologically relevant measures. This new mechanism acts as a independent ”system test” for computer vision algorithms, as errors in different parts of a pipeline can propagate in unpredictable ways [8, 48]. While these newly proposed metrics are not meant to replace tra- ditional task-specific benchmarking, they acts as a bridge to allow biologists to directly evaluate whether computer vision solutions are sufficient for application, and directly apply architectures appropriately. Another feature of application-driven ML is the avail- ability of domain-specific knowledge when designing com- puter vision solutions [50]. We incorporate this concept throughout the design and benchmarking of the CHIRP dataset, by providing color ring segmentation and metadata on the probable birds to constraint the re-id problem. This added layer of complexity encourages future computer vi- sion solutions to make use of the constraints of the study system, as we demonstrate with the CORVID, which iden- tifies individuals purely from detecting the presence of color rings, contrary to traditional re-id approaches. The CORVID pipeline is more flexible than traditional approaches as it does not rely on a gallery to be matched with, and only relies on a list of possible color combina- tions. However, benchmarks on CHIRP showed that the accuracy of the framework depends on the constraint where only limited birds can be present in a video (Table 1), mak- ing the method unscalable to larger bird populations. Future work can combine image based methods to compute simi- larities of appearances as demonstrated in other bird species [20, 62]. This will be important for birds as many birds change appearance over time or seasons. Now, we discuss some limitations with the CHIRP dataset. Firstly, in contrary to the ChimpACT dataset [36], annotations presented in the current dataset are done in dif- ferent data subsets. This stems from the annotation strat- egy we employ, focusing our manual annotation efforts to diverse frames instead of video sequences like in Chim- pACT [36]. To provide a solution to this problem, we pro- vide model-annotated 2D keypoints and segmentation on the re-id and action recognition datasets. Next, compared to other action recognition datasets, the number of annotations and behavioral classes provided in the dataset is limited, and is highly skewed towards feed- ing behavior (Figure 3). This is primarily a reflection of the distribution of behaviors that jays perform on the feeder, as they primarily spend their time feeding. The specific be- haviors required depends on the final use case, however, quantifying feeding and individual presence/ co-presence is valuable for long-term behavioral monitoring, in relation to food acquisition and understanding the evolution of social behaviors [15, 27, 39]. In addition, other behaviors like vig- ilance on the feeder (head held upright to identify predators [25, 26]) can be reliably extracted from body postures if 2D keypoint estimates are accurate, and thus the types of behav- iors that can be extracted using this dataset is not limited by the action recognition dataset. Finally, the CHIRP dataset only includes data from a sin- gle species and study system, making it hard to evaluate whether developed algorithms can be generalized. How- ever, CHIRP is a first of its kind dataset of a wild long- term bird population, which is limited by the years of effort that is required to establish and collect. We hope that the approach for preparing the CHIRP dataset, together with the novel application-specific benchmarking can act as a blueprint for future work, on how to design datasets care- fully to encourage computer vision algorithms to be read- ily applied to the application domain. With well designed datasets and algorithms, computer vision can potentially revolutionize individual level behavior monitoring for the study of animal behavior, conservation and beyond. 8 7 Acknowledgments This work is funded by the Deutsche Forschungsge- meinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy—EXC 2117—422037984, DFG project 15990824, DFG Heisenberg Grant no. GR 4650/21 and DFG project Grant no. FP589/20. MHW is supported by an Oxford Brookes University Emerging Leaders Research Fellowship. We thank Francesca Frisoni and Jyotsna Bellary for doing additional annotations for the application specific benchmark. We thank all the re- searchers and field workers who worked in the Luondua Boreal Field Station over the years for their contributions to the long-term dataset. References [1] Gustavo Alarc ́ on-Nieto, Jacob M Graving, James A Klarevas-Irby,Adriana A Maldonado-Chaparro,Inge Mueller, and Damien R Farine. An automated barcode track- ing system for behavioural studies in birds. Methods in Ecol- ogy and Evolution, 9(6):1536–1547, 2018. 5 [2] Guy Q.A. Anderson and Rhys E. Green. The value of ringing for bird conservation. Ringing & Migration, 24(3):205–212, 2009. 2 [3] Th ́ eo Ardoin and C ́ edric Sueur. Automatic Identification of Stone-Handling Behaviour in Japanese Macaques Using LabGym Artificial Intelligence, 2023.arXiv:2310.07812 [cs]. 1 [4] Otto Brookes, Majid Mirmehdi, Colleen Stephens, Samuel Angedakin, Katherine Corogenes, Dervla Dowd, Paula Dieguez, Thurston C. Hicks, Sorrel Jones, Kevin Lee, Vera Leinert, Juan Lapuente, Maureen S. McCarthy, Amelia Meier, Mizuki Murai, Emmanuelle Normand, Virginie Vergnes, Erin G. Wessling, Roman M. Wittig, Kevin Langer- graber, Nuria Maldonado, Xinyu Yang, Klaus Zuberb ̈ uhler, Christophe Boesch, Mimi Arandjelovic, Hjalmar K ̈ uhl, and Tilo Burghardt. Panaf20k: A large video dataset for wild ape detection and behaviour recognition. International Journal of Computer Vision, 132(8):3086–3102, 2024. 2 [5] Otto Brookes, Maksim Kukushkin, Majid Mirmehdi, Colleen Stephens, Paula Dieguez, Thurston C. Hicks, Sor- rel Jones, Kevin Lee, Maureen S. McCarthy, Amelia Meier, Emmanuelle Normand, Erin G. Wessling, Roman M. Wit- tig, Kevin Langergraber, Klaus Zuberb ̈ uhler, Lukas Boesch, Thomas Schmid, Mimi Arandjelovic, Hjalmar K ̈ uhl, and Tilo Burghardt. The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour Recogni- tion. pages 5433–5443, 2025. 2 [6] Mikel Brostr ̈ om. BoxMOT: pluggable SOTA tracking mod- ules for object detection, segmentation and pose estimation models. 7 [7] B Calvo and RW Furness. A review of the use and the effects of marks and devices on birds. Ringing & Migration, 13(3): 129–151, 1992. 2 [8] Alex Hoi Hang Chan, Otto Brookes, Urs Waldmann, Hemal Naik, Iain D. Couzin, Majid Mirmehdi, No ̈ el Adiko Houa, Emmanuelle Normand, Christophe Boesch, Lukas Boesch, Mimi Arandjelovic, Hjalmar K ̈ uhl, Tilo Burghardt, and Fu- mihiro Kano. Towards Application-Specific Evaluation of Vision Models: Case Studies in Ecology and Biology, 2025. 2, 7, 8 [9] Hoi Hang Chan, Prasetia Putra, Harald Schupp, Johanna K ̈ ochling, Jana Straßheim, Britta Renner, Julia Schroeder, William D Pearse, Shinichi Nakagawa, Terry Burke, Michael Griesser, Andrea Meltzer, Saverio Lubrano, and Fumihiro Kano. Yolo-behaviour: A simple, flexible framework to au- tomatically quantify animal behaviours from videos. Meth- ods in Ecology and Evolution, 16(4):760–774, 2025. 1 [10] Jun Chen, Ming Hu, Darren J Coker, Michael L Berumen, Blair Costelloe, Sara Beery, Anna Rohrbach, and Mohamed Elhoseiny. Mammalnet: A large-scale video benchmark for mammal recognition and behavior understanding. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 13052–13061, 2023. 2 [11] Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexan- der Kirillov, and Rohit Girdhar.Masked-attention Mask Transformer for Universal Image Segmentation, 2022. arXiv:2112.01527 [cs]. 6 [12] Michael Chimento, Alex Hoi Hang Chan, Lucy M Aplin, and Fumihiro Kano. Peering into the world of wild passer- ines with 3d-socs: Synchronized video capture and posture estimation. Methods in Ecology and Evolution, 2025. 1 [13] James D. Crall, Nick Gravish, Andrew M. Mountcastle, and Stacey A. Combes. BEEtag: A Low-Cost, Image-Based Tracking System for the Study of Animal Behavior and Lo- comotion. PLOS ONE, 10(9):e0136487, 2015. Publisher: Public Library of Science. 5 [14] Antica Culina, Frank Adriaensen, Liam D Bailey, Mal- colm D Burgess, Anne Charmantier, Ella F Cole, Tapio Eeva, Erik Matthysen, Chlo ́ e R Nater, Ben C Sheldon, et al. Con- necting the data landscape of long-term ecological studies: The spi-birds data hub. Journal of Animal Ecology, 90(9): 2147–2160, 2021. 5 [15] Filipe CR Cunha and Michael Griesser. Who do you trust? wild birds use social knowledge to avoid being deceived. Sci- ence Advances, 7(22):eaba2862, 2021. 8 [16] Nisarg Desai, Praneet Bala, Rebecca Richardson, Jessica Raper, Jan Zimmermann, and Benjamin Hayden. OpenApe- Pose, a database of annotated ape photographs for pose esti- mation. eLife, 12:RP86873, 2023. Publisher: eLife Sciences Publications, Ltd. 1, 2 [17] Isla Duporge, Maksim Kholiavchenko, Roi Harel, Scott Wolf, Daniel I Rubenstein, Margaret C Crofoot, Tanya Berger-Wolf, Stephen J Lee, Julie Barreau, Jenna Kline, et al. Baboonland dataset: Tracking primates in the wild and au- tomating behaviour recognition from drone videos: I. du- porge et al. International Journal of Computer Vision, pages 1–12, 2025. 2 [18] Jan Ekman and Michael Griesser. Siberian jays: Delayed dispersal in the absence of cooperative breeding. In Coop- erative Breeding in Vertebrates: Studies of Ecology, Evolu- tion, and Behavior, pages 6–18. Cambridge University Press, Cambridge, 2016. 3 [19] European colour-ring Birding. European colour-ring birding. 9 https://cr-birding.org/, 1995–. Accessed: 2026- 03-03. 5 [20] Andr ́ e C Ferreira, Liliana R Silva, Francesco Renna, Hanja B Brandl, Julien P Renoult, Damien R Farine, Rita Covas, and Claire Doutrelant. Deep learning-based methods for indi- vidual recognition in small birds. Methods in Ecology and Evolution, 11(9):1072–1085, 2020. 1, 3, 5, 8 [21] Olivier Friard and Marco Gamba. BORIS: a free, versatile open-source event-logging software for video/audio coding and live observations. Methods in Ecology and Evolution, 7 (11):1325–1330, 2016. 4 [22] Michael Fuchs, Emilie Genty, Klaus Zuberb ̈ uhler, and Paul Cotofrei. ASBAR: an Animal Skeleton-Based Action Recognition framework. Recognizing great ape behaviors in the wild using pose estimation with domain adaptation. eLife, 13, 2024. Publisher: eLife Sciences Publications Lim- ited. 1 [23] Valentin Gabeff, Haozhe Qi, Brendan Flaherty, Gencer Sum- bul, Alexander Mathis, and Devis Tuia. MammAlps: A Multi-view Video Behavior Monitoring Dataset of Wild Mammals in the Swiss Alps. pages 13854–13864, 2025. 2 [24] Jacob M Graving, Daniel Chae, Hemal Naik, Liang Li, Ben- jamin Koger, Blair R Costelloe, and Iain D Couzin. Deep- posekit, a software toolkit for fast and robust animal pose estimation using deep learning. Elife, 8:e47994, 2019. 1 [25] Michael Griesser. Nepotistic vigilance behavior in Siberian jay parents. Behavioral Ecology, 14(2):246–250, 2003. 8 [26] Michael Griesser and Magdalena Nystrand. Vigilance and predation of a forest-living bird species depend on large- scale habitat structure. Behavioral Ecology, 20(4):709–715, 2009. 8 [27] Michael Griesser, Peter Halvarsson, Szymon M. Drobniak, and Carles Vil ` a. Fine-scale kin recognition in the absence of social familiarity in the Siberian jay, a monogamous bird species. Molecular Ecology, 24(22):5726–5738, 2015. 4, 8 [28] Michael Griesser, Szymon M Drobniak, Shinichi Nakagawa, and Carlos A Botero. Family living sets the stage for co- operative breeding and ecological resilience in birds. PLoS biology, 15(6):e2000483, 2017. 4 [29] Glenn Jocher, Ayush Chaurasia, and Jing Qiu. Ultralytics YOLO, 2023. 4 [30] Daniel Joska, Liam Clark, Naoya Muramatsu, Ricardo Jericevich, Fred Nicolls, Alexander Mathis, Mackenzie W Mathis, and Amir Patel. Acinoset: a 3d pose estimation dataset and baseline models for cheetahs in the wild. In 2021 IEEE international conference on robotics and automation (ICRA), pages 13901–13908. IEEE, 2021. 1, 2 [31] Maksim Kholiavchenko, Jenna Kline, Michelle Ramirez, Sam Stevens, Alec Sheets, Reshma Babu, Namrata Banerji, Elizabeth Campolongo, Matthew Thompson, Nina Van Tiel, Jackson Miliko, Eduardo Bessa, Isla Duporge, Tanya Berger- Wolf, Daniel Rubenstein, and Charles Stewart. KABR: In- Situ Dataset for Kenyan Animal Behavior Recognition from Drone Videos. In 2024 IEEE/CVF Winter Conference on Ap- plications of Computer Vision Workshops (WACVW), pages 31–40, Waikoloa, HI, USA, 2024. IEEE. 2 [32] Benjamin Koger, Edward Hurme, Blair R. Costelloe, M. Teague O’Mara, Martin Wikelski, Roland Kays, and Dina K. N. Dechmann. An automated approach for count- ing groups of flying animals applied to one of the world’s largest bat colonies. Ecosphere, 14(6):e4590, 2023. 1 [33] Hjalmar S. K ̈ uhl and Tilo Burghardt. Animal biometrics: quantifying and detecting phenotypic appearance. Trends in Ecology & Evolution, 28(7):432–441, 2013. Publisher: El- sevier. 1 [34] Ci Li, Ylva Mellbin, Johanna Krogager, Senya Polikovsky, Martin Holmberg, Nima Ghorbani, Michael J. Black, Hedvig Kjellstr ̈ om, Silvia Zuffi, and Elin Hernlund. The Poses for Equine Research Dataset (PFERD). Scientific Data, 11(1): 497, 2024. Publisher: Nature Publishing Group. 2 [35] Dan Liu, Jin Hou, Shaoli Huang, Jing Liu, Yuxin He, Bochuan Zheng, Jifeng Ning, and Jingdong Zhang. Lote- animal: A long time-span dataset for endangered animal be- havior understanding. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision, pages 20064– 20075, 2023. 2 [36] Xiaoxuan Ma, Stephan Kaufhold, Jiajun Su, Wentao Zhu, Jack Terwilliger, Andres Meza, Yixin Zhu, Federico Rossano, and Yizhou Wang.Chimpact: A longitudinal dataset for understanding chimpanzee behaviors. Advances in Neural Information Processing Systems, 36:27501–27531, 2023. 2, 3, 8 [37] Xiaoxuan Ma, Yutang Lin, Yuan Xu, Stephan P. Kaufhold, Jack Terwilliger, Andres Meza, Yixin Zhu, Federico Rossano, and Yizhou Wang.AlphaChimp:Track- ing and Behavior Recognition of Chimpanzees, 2024. arXiv:2410.17136 [cs]. 3 [38] Jesse D Marshall, Ugne Klibaite, Amanda Gellis, Diego E Aldarondo, Bence P ̈ Olveczky, and Timothy W Dunn. The pair-r24m dataset for multi-animal 3d pose estimation. bioRxiv, pages 2021–11, 2021. 2 [39] Andrea Meltzer. The ecological and social drivers of co- operation: Experimental field studies in a group-living bird. 2025. 8 [40] Luke Meyers, Josu ́ e Rodr ́ ıguez Cordero, Carlos Corrada Bravo, Fanfan Noel, Jos ́ e Agosto-Rivera, Tugrul Giray, and R ́ emi M ́ egret. Towards automatic honey bee flower-patch assays with paint marking re-identification. arXiv preprint arXiv:2311.07407, 2023. 5 [41] MMAction2-Contributors. OpenMMLab’s Next Generation Video Understanding Toolbox and Benchmark, 2020. 6 [42] MMPose-Contributors. OpenMMLab Pose Estimation Tool- box and Benchmark, 2020. 6 [43] M ́ at ́ e Nagy, G ́ abor V ́ as ́ arhelyi, Benjamin Pettit, Isabella Roberts-Mariani, Tam ́ as Vicsek, and Dora Biro. Context- dependent hierarchies in pigeons. Proceedings of the Na- tional Academy of Sciences, 110(32):13049–13054, 2013. 5 [44] M ́ at ́ e Nagy, Jacob D Davidson, G ́ abor V ́ as ́ arhelyi, D ́ aniel ́ Abel, Enik ̋ o Kubinyi, Ahmed El Hady, and Tam ́ as Vicsek. Long-term tracking of social structure in groups of rats. Sci- entific Reports, 14(1):22857, 2024. 5 [45] Hemal Naik, Alex Hoi Hang Chan, Junran Yang, Mathilde Delacoux, Iain D Couzin, Fumihiro Kano, and M ́ at ́ e Nagy. 3d-pop-an automated annotation approach to facilitate mark- erless 2d-3d tracking of freely moving birds with marker- 10 based motion capture. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition, pages 21274–21284, 2023. 2, 3 [46] Hemal Naik, Junran Yang, Dipin Das, Margaret Cro- foot, Akanksha Rathore, and Vivek Hari Sridhar. Buck- tales: A multi-uav dataset for multi-object tracking and re- identification of wild antelopes. Advances in Neural Infor- mation Processing Systems, 37:81992–82009, 2024. 3 [47] Xun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni, Si Yong Yeo, and Jun Liu. Animal kingdom: A large and diverse dataset for animal behavior understanding. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 19023–19034, 2022. 2 [48] Omiros Pantazis, Peggy Bevan, Holly Pringle, Guil- herme Braga Ferreira, Daniel J Ingram, Emily Madsen, Liam Thomas, Dol Raj Thanet, Thakur Silwal, Santosh Rayama- jhi, et al. Deep learning-based ecological analysis of cam- era trap images is impacted by training data quality and size. arXiv preprint arXiv:2408.14348, 2024. 2, 8 [49] Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R ̈ adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao- Yuan Wu, Ross Girshick, Piotr Doll ́ ar, and Christoph Feicht- enhofer. SAM 2: Segment Anything in Images and Videos, 2024. arXiv:2408.00714 [cs]. 5 [50] David Rolnick, Alan Aspuru-Guzik, Sara Beery, Bistra Dilk- ina, Priya L. Donti, Marzyeh Ghassemi, Hannah Kerner, Claire Monteleoni, Esther Rolf, Milind Tambe, and Adam White. Position: Application-Driven Innovation in Machine Learning. In Proceedings of the 41st International Con- ference on Machine Learning, pages 42707–42718. PMLR, 2024. ISSN: 2640-3498. 7, 8 [51] Gabriel Santiago-Plaza, Luke Meyers, Andrea Gomez- Jaime, Rafael Mel ́ endez-Rıos, Fanfan Noel, Jose Agosto, Tugrul Giray, Josu ́ e Rodrıguez-Cordero, and R ́ emi M ́ egret. Identification of honeybees with paint codes using convolu- tional neural networks. Proceedings of the 19th International Joint Conference on Computer Vision, Imaging and Com- puter Graphics Theory and Applications, 772:779, 2024. 5 [52] Stefan Schneider, Graham W. Taylor, Stefan Linquist, and Stefan C. Kremer. Past, present and future approaches us- ing computer vision for animal re-identification from cam- era trap data. Methods in Ecology and Evolution, 10(4):461– 470, 2019. 1 [53] Andrew Sih, Maud CO Ferrari, and David J Harris. Evolu- tion and behavioural responses to human-induced rapid envi- ronmental change. Evolutionary applications, 4(2):367–387, 2011. 1 [54] Devis Tuia, Benjamin Kellenberger, Sara Beery, Blair R Costelloe, Silvia Zuffi, Benjamin Risse, Alexander Mathis, Mackenzie W Mathis,Frank van Langevelde,Tilo Burghardt, et al.Perspectives in machine learning for wildlife conservation. Nature communications, 13(1):792, 2022. 2 [55] Ulla Tuomainen and Ulrika Candolin. Behavioural responses to human-induced environmental change. Biological Re- views, 86(3):640–657, 2011. 1 [56] Yana van de Sande, Wim Pouw, and Lara M. Southern. Au- tomated Recognition of Grooming Behavior in Wild Chim- panzees. Proceedings of the Annual Meeting of the Cognitive Science Society, 46(0), 2024. 1 [57] Urs Waldmann, Alex Hoi Hang Chan, Hemal Naik, M ́ at ́ e Nagy, Iain D. Couzin, Oliver Deussen, Bastian Goldluecke, and Fumihiro Kano. 3D-MuPPET: 3D Multi-Pigeon Pose Estimation and Tracking. International Journal of Computer Vision, 2024. 1 [58] Fasheng Wang, Ping Cao, Fu Li, Xing Wang, Bing He, and Fuming Sun. WATB: Wild Animal Tracking Benchmark. International Journal of Computer Vision, 131(4):899–917, 2023. 2 [59] Miyako H Warrington, David N Fisher, Jan Komdeur, Na- talie Pilakouta, and Michael Griesser. Stronger together? a framework for studying population resilience to climate change impacts via social shielding. 2024. 4 [60] Charlotte Wiltshire, James Lewis-Cheetham, Viola Kome- dov ́ a, Tetsuro Matsuzawa, Kirsty E. Graham, and Cather- ine Hobaiter. DeepWild: Application of the pose estima- tion tool DeepLabCut for behaviour tracking in wild chim- panzees and bonobos. Journal of Animal Ecology, 92(8): 1560–1574, 2023. 1 [61] Scott W. Wolf, Dee M. Ruttenberg, Daniel Y. Knapp, An- drew E. Webb, Ian M. Traniello, Grace C. McKenzie-Smith, Sophie A. Leheny, Joshua W. Shaevitz, and Sarah D. Kocher. NAPS: Integrating pose estimation and tag-based track- ing. Methods in Ecology and Evolution, 14(10):2541–2548, 2023. 5 [62] Shiting Xiao, Yufu Wang, Ammon Perkes, Bernd Pfrommer, Marc Schmidt, Kostas Daniilidis, and Marc Badger. Multi- view tracking, re-id, and social network analysis of a flock of visually similar birds in an outdoor aviary. International Journal of Computer Vision, 131(6):1532–1549, 2023. 3, 5, 8 [63] Shaokai Ye, Anastasiia Filippova, Jessy Lauer, Steffen Schneider, Maxime Vidal, Tian Qiu, Alexander Mathis, and Mackenzie Weygandt Mathis. Superanimal pretrained pose estimation models for behavioral analysis. Nature Commu- nications, 15(1):5165, 2024. 1 [64] Hang Yu, Yufei Xu, Jing Zhang, Wei Zhao, Ziyu Guan, and Dacheng Tao. Ap-10k: A benchmark for animal pose es- timation in the wild. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. 2 [65] Libo Zhang, Junyuan Gao, Zhen Xiao, and Heng Fan. Ani- malTrack: A Benchmark for Multi-Animal Tracking in the Wild. International Journal of Computer Vision, 131(2): 496–513, 2023. 2 [66] Vojt ˇ ech ˇ Cerm ́ ak, Lukas Picek, Luk ́ a ˇ s Adam, and Kostas Pa- pafitsoros. WildlifeDatasets: An Open-Source Toolkit for Animal Re-Identification. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 5953–5963, 2024. 2, 4, 6 11