Paper deep dive
YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition
Serdar Yildiz, Abbas MemiĆ, SongĂŒl Varli
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/19/2026, 4:07:54 AM
Summary
The paper introduces YILDIZ-VPR, a novel dataset for Visual Place Recognition (VPR) designed to address the need for dense, pedestrian-level visual data under diverse environmental conditions. Collected via repeated walking traversals on the Davutpasa campus of Yildiz Technical University using a GoPro 9 camera, the dataset includes 370,464 training images and over 4,000 query images (daytime and nighttime). It features synchronized GPS, gyroscope, speed, and temperature metadata, offering high spatial density and long-term visual variability to support research in image-based and temporal VPR.
Entities (11)
Relation Signals (9)
YILDIZ-VPR â collectedat â Davutpasa Campus
confidence 95% · The dataset was recorded on the Davutpasa campus of Yildiz Technical University
YILDIZ-VPR â supportstask â Visual Place Recognition
confidence 95% · YILDIZ-VPR provides a useful resource for studying image-based and temporal visual place recognition
YILDIZ-VPR â createdby â Serdar Yıldız
confidence 90% · Serdar Yıldız... In this paper, we introduce YILDIZ-VPR
YILDIZ-VPR â createdby â Abbas MemiĆ
confidence 90% · Abbas MemiĆ... In this paper, we introduce YILDIZ-VPR
YILDIZ-VPR â createdby â SongĂŒl Varlı
confidence 90% · SongĂŒl Varlı... In this paper, we introduce YILDIZ-VPR
YILDIZ-VPR â hasmetadata â GPS
confidence 90% · In addition to GPS coordinates, the dataset also includes auxiliary sensor information
YILDIZ-VPR â useshardware â GoPro 9
confidence 90% · Each video was recorded with a GoPro 9 camera
YILDIZ-VPR â iscomparedto â Oxford RobotCar
confidence 80% · Later datasets, including... Oxford RobotCar... further extended the field
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Visual Place Recognition (VPR) aims to recognize the location of a query image by comparing it with a set of geo-referenced images. Although many datasets have been proposed for VPR, collecting dense and diverse visual data from pedestrian-level viewpoints is still an important need. In this paper, we introduce YILDIZ-VPR, a visual geo-localization dataset collected through repeated walking traversals on the Davutpasa campus of Yildiz Technical University. The dataset includes outdoor scenes captured at different times of day, seasons, and weather conditions. It contains a wide range of visual content, including historical buildings, modern structures, roads, green areas, and wooded regions. Each video was recorded with a GoPro 9 camera and synchronized with GPS sensor data to provide location labels for the extracted frames. In addition to GPS coordinates, the dataset also includes auxiliary sensor information such as gyroscope, speed, and temperature data. With its dense coverage and long-term visual variability, YILDIZ-VPR provides a useful resource for studying image-based and temporal visual place recognition under realistic outdoor conditions.
Tags
Links
- Source: https://arxiv.org/abs/2608.17033v1
- Canonical: https://arxiv.org/abs/2608.17033v1
Trouble viewing inline? Open PDF directly â
Full Text
16,187 characters extracted from source content.
Expand or collapse full text
YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition Serdar Yıldız Affiliation: Department of Computer Engineering Affiliation: Yildiz Technical University, Istanbul, TĂŒrkiye Affiliation: BILGEM, TUBITAK, Gebze, Kocaeli, TĂŒrkiye Email: serdar.yildiz@std.yildiz.edu.tr Abbas MemiĆ Affiliation: Department of Computer Engineering Affiliation: Istanbul University, Istanbul, TĂŒrkiye Email: abbas.memis@istanbul.edu.tr SongĂŒl Varlı Affiliation: Department of Computer Engineering Affiliation: Yildiz Technical University, Istanbul, TĂŒrkiye Email: svarli@yildiz.edu.tr Abstract Visual Place Recognition (VPR) aims to recognize the location of a query image by comparing it with a set of geo-referenced images. Although many datasets have been proposed for VPR, collecting dense and diverse visual data from pedestrian-level viewpoints is still an important need. In this paper, we introduce YILDIZ-VPR, a visual geo-localization dataset collected through repeated walking traversals on the Davutpasa campus of Yildiz Technical University. The dataset includes outdoor scenes captured at different times of day, seasons, and weather conditions. It contains a wide range of visual content, including historical buildings, modern structures, roads, green areas, and wooded regions. Each video was recorded with a GoPro 9 camera and synchronized with GPS sensor data to provide location labels for the extracted frames. In addition to GPS coordinates, the dataset also includes auxiliary sensor information such as gyroscope, speed, and temperature data. With its dense coverage and long-term visual variability, YILDIZ-VPR provides a useful resource for studying image-based and temporal visual place recognition under realistic outdoor conditions. The YILDIZ-VPR dataset is available at https://w.github.com/⊠1 Introduction Figure 1: Sample images from the YILDIZ-VPR dataset. Visual Place Recognition (VPR) aims to identify the location of a visual observation by matching it against a set of previously collected geo-referenced images. It is commonly formulated as an image retrieval problem, where a query image is represented by a visual descriptor and compared with gallery images whose locations are known 9; 16. This formulation has made VPR an important component of visual localization systems, with applications in autonomous navigation, mobile robotics, augmented reality, and location-aware visual search. Despite its practical importance, robust place recognition remains challenging in real-world environments. A single place may exhibit substantially different appearances depending on illumination, weather, season, time of day, viewpoint, shadows, and dynamic objects. These factors become particularly critical in dense localization scenarios, where visually similar locations may be only a few meters apart. Therefore, progress in VPR depends not only on stronger visual representations but also on datasets that reflect the spatial density and environmental variability observed in practical navigation settings. GPS-based localization is widely used in modern positioning and navigation systems 7. However, GPS measurements may be noisy or unreliable in outdoor areas surrounded by buildings, trees, narrow paths, or other sources of signal degradation. In such cases, visual information can provide a complementary cue for localization. Rather than replacing GPS, VPR can support location estimation by using the visual structure of the environment, especially when accurate or stable coordinate measurements are difficult to obtain. Several public datasets have contributed significantly to VPR research by providing large-scale, long-term, or cross-condition visual data. Early and widely used datasets such as St Lucia, Eynsham, San Francisco, Nordland, Tokyo 24/7, and Pitts250k introduced different forms of geographical coverage, viewpoint diversity, temporal variation, and illumination change 10; 6; 4; 11; 12; 13. Later datasets, including SPED, Oxford RobotCar, Mapillary SLS, SVOX, AmsterTime, GSV-CITIES, and San Francisco XL, further extended the field with long-term recordings, large-scale street-level imagery, historical views, multi-city coverage, and challenging environmental changes 5; 8; 14; 3; 15; 1; 2. These datasets have shaped the evaluation of VPR methods and remain highly valuable for the community. However, many existing datasets are collected using vehicle-mounted cameras, trains, surveillance cameras, or street-view platforms. While these acquisition strategies are effective for road-based and city-scale localization, they do not fully represent pedestrian-level visual perception. Road-centric data may provide limited access to walking paths, open spaces, building surroundings, green areas, and transitions between different outdoor regions. For applications involving pedestrians, wearable cameras, or mobile robots operating in human-scale environments, dense data collected from walking trajectories is therefore an important complementary resource. Another challenge is the joint availability of dense spatial coverage and repeated observations under diverse environmental conditions. A dataset may cover a large geographical region, but this does not necessarily mean that the same locations are observed across different times of day, seasons, weather conditions, and viewpoints. Repeated observations of the same environment provide a more suitable basis for studying how visual place representations change over time. Such data are particularly useful for research on day-night matching, long-term localization, temporal VPR, and robustness under environmental variation. To address these needs, we introduce YILDIZ-VPR, a dense visual place recognition dataset collected through repeated walking traversals. The dataset was recorded on the Davutpasa campus of Yildiz Technical University, which contains a diverse set of outdoor scenes, including historical buildings, modern structures, roads, green areas, open spaces, and wooded regions. This environment provides a compact yet visually rich setting, where natural and man-made structures coexist within a relatively dense spatial area. 2 YILDIZ-VPR Dataset YILDIZ-VPR was designed to provide dense, pedestrian-level visual data for visual place recognition research. The dataset focuses on outdoor place recognition under realistic environmental changes, rather than controlled or single-session image matching. For this purpose, repeated walking traversals were recorded in the same environment across different times of day, weather conditions, and seasonal periods. This design makes the dataset suitable for studying both spatially dense localization and long-term visual variation. Figure 2: Timeline and sample diversity of the YILDIZ-VPR dataset. 2.1 Collection Environment The dataset was collected on the Davutpasa campus of Yildiz Technical University. This location was selected because it contains a diverse mixture of outdoor scene types within a compact geographical area. The campus includes historical buildings, modern architectural structures, roads, open spaces, grass-covered areas, and wooded regions. These elements create a visually rich environment where natural and man-made structures appear together. Indoor areas were not included in the dataset. This decision was made to keep the focus on outdoor visual place recognition, where environmental changes such as lighting, weather, shadows, and seasonal appearance differences play an important role. Since many outdoor locations on the campus are visually close to each other, the dataset also contains challenging cases in which nearby places may share similar visual patterns. 2.2 Acquisition Protocol YILDIZ-VPR was recorded using a GoPro 9 camera during walking traversals. The videos were captured at 4K resolution and 30 frames per second. Unlike vehicle-mounted or street-view datasets, the camera viewpoint follows a pedestrian-level trajectory. This provides visual observations that are closer to how a person, wearable camera, or mobile robot perceives an outdoor environment. The same areas were visited multiple times to capture the visual appearance of locations under different conditions. Some traversals were performed close to each other in time to increase viewpoint diversity, while others were recorded across different periods to capture changes in illumination, season, and weather. As a result, the dataset contains both fine-grained spatial coverage and repeated observations of the same environment under changing conditions. 2.3 GPS-based Annotation Each recorded video is paired with GPS sensor data. The image frames were synchronized with the GPS measurements to assign geographical coordinates to the visual samples. Since GPS accuracy can be affected by buildings, trees, weather conditions, and other environmental factors, low-quality GPS measurements were filtered during dataset preparation. Only frames associated with GPS accuracy values below 5 meters were included in the dataset. After filtering, the dataset provides location annotations with an average GPS accuracy of approximately 1.5 meters. The GPS signal was recorded approximately every 122 milliseconds, corresponding to a frequency of about 8â9 Hz. The videos were sampled in synchronization with the GPS stream at approximately 4 Hz. Depending on walking speed, this sampling strategy corresponds to roughly four images per meter, resulting in dense visual coverage of the traversed routes. 2.4 Dataset Organization YILDIZ-VPR is organized into three main parts: training, daytime test, and nighttime test. This structure was designed to support the analysis of environmental and temporal changes in visual place recognition. The training set contains densely sampled frames from the main walking traversals. The daytime test set includes query images captured during daytime conditions, while the nighttime test set contains query images recorded at night. Figure 3: Spatial distribution of the YILDIZ-VPR dataset samples. The training frames were sampled from the recorded videos at regular intervals. In contrast, the query images in the test sets were manually selected to reduce redundancy between consecutive frames. This selection strategy increases the diversity of the query images and avoids constructing test sets from nearly identical adjacent frames. The final dataset contains 370,464 training images, 2,377 daytime query images, and 1,902 nighttime query images. The training set can be used as the gallery set for image retrieval-based VPR experiments, while the daytime and nighttime query sets can be used to study localization under different illumination conditions. 2.5 Metadata In addition to RGB image data and GPS coordinates, YILDIZ-VPR includes auxiliary sensor metadata. These metadata include sensor temperature, speed, and three-axis gyroscope measurements. Such information can provide additional context about the camera motion and acquisition conditions. Although the dataset can be directly used for image-based VPR, the availability of video sequences and synchronized sensor data also makes it suitable for future temporal place recognition studies. For example, sequential models may use consecutive frames, motion cues, or sensor information to improve localization robustness under challenging visual conditions. Therefore, YILDIZ-VPR is not limited to single-image retrieval settings and can also support future research on temporal and sensor-aware visual localization. 3 Conclusion In this work, we introduced YILDIZ-VPR, a densely sampled visual place recognition dataset collected through repeated pedestrian-level traversals of the Davutpasa campus of Yildiz Technical University under different times of day, seasons, weather conditions, and illumination settings. The dataset contains 370,464 training images, 2,377 daytime query images, and 1,902 nighttime query images, covering diverse outdoor scenes such as historical and modern buildings, roads, walking paths, open spaces, green areas, and wooded regions. Each image is associated with GPS coordinates. By combining dense spatial coverage, repeated observations of the same environment, pedestrian-level viewpoints, and substantial long-term appearance variation, YILDIZ-VPR offers a complementary resource for investigating image-based, sequential, day-to-night, and multimodal visual place recognition. Future work will establish comprehensive evaluation protocols and benchmark representative VPR methods on the dataset. References Ali-bey et al. (2022) A. Ali-bey, B. Chaib-draa, and P. GiguĂšre GSV-cities: toward appropriate supervised visual place recognition. Neurocomputing 513, p. 194â203. Cited by: §1. Berton et al. (2022) G. Berton, C. Masone, and B. Caputo Rethinking visual geo-localization for large-scale applications. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 4878â4888. Cited by: §1. Berton et al. (2021) G. M. Berton, V. Paolicelli, C. Masone, and B. Caputo Adaptive-attentive geolocalization from few queries: a hybrid approach. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p. 2918â2927. Cited by: §1. Chen et al. (2011) D. M. Chen, G. Baatz, K. Köser, S. S. Tsai, R. Vedantham, T. PylvĂ€nĂ€inen, K. Roimela, X. Chen, J. Bach, M. Pollefeys, et al. City-scale landmark identification on mobile devices. In CVPR 2011, p. 737â744. Cited by: §1. Chen et al. (2017) Z. Chen, A. Jacobson, N. SĂŒnderhauf, B. Upcroft, L. Liu, C. Shen, I. Reid, and M. Milford Deep learning features at scale for visual place recognition. In 2017 IEEE international conference on robotics and automation (ICRA), p. 3223â3230. Cited by: §1. Cummins and Newman (2009) M. Cummins and P. Newman Highly scalable appearance-only slam-fab-map 2.0.. In Robotics: Science and systems, Vol. 5, p. 17. Cited by: §1. Drawil et al. (2012) N. M. Drawil, H. M. Amar, and O. A. Basir GPS localization accuracy classification: a context-based approach. IEEE Transactions on Intelligent Transportation Systems 14 (1), p. 262â273. Cited by: §1. Maddern et al. (2017) W. Maddern, G. Pascoe, C. Linegar, and P. Newman 1 year, 1000 km: the oxford robotcar dataset. The International Journal of Robotics Research 36 (1), p. 3â15. Cited by: §1. Masone and Caputo (2021) C. Masone and B. Caputo A survey on deep visual place recognition. IEEE Access 9, p. 19516â19547. Cited by: §1. Milford and Wyeth (2008) M. J. Milford and G. F. Wyeth Mapping a suburb with a single camera using a biologically inspired slam system. IEEE Transactions on Robotics 24 (5), p. 1038â1053. Cited by: §1. SĂŒnderhauf et al. (2013) N. SĂŒnderhauf, P. Neubert, and P. Protzel Are we there yet? challenging seqslam on a 3000 km journey across all four seasons. In Proc. of workshop on long-term autonomy, IEEE international conference on robotics and automation (ICRA), p. 2013. Cited by: §1. Torii et al. (2015) A. Torii, R. Arandjelovic, J. Sivic, M. Okutomi, and T. Pajdla 24/7 place recognition by view synthesis. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 1808â1817. Cited by: §1. Torii et al. (2013) A. Torii, J. Sivic, T. Pajdla, and M. Okutomi Visual place recognition with repetitive structures. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 883â890. Cited by: §1. Warburg et al. (2020) F. Warburg, S. Hauberg, M. Lopez-Antequera, P. Gargallo, Y. Kuang, and J. Civera Mapillary street-level sequences: a dataset for lifelong place recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 2626â2635. Cited by: §1. Yildiz et al. (2022) B. Yildiz, S. Khademi, R. M. Siebes, and J. Van Gemert Amstertime: a visual place recognition benchmark dataset for severe domain shift. In 2022 26th International Conference on Pattern Recognition (ICPR), p. 2749â2755. Cited by: §1. Zhang et al. (2021) X. Zhang, L. Wang, and Y. Su Visual place recognition: a survey from deep learning perspective. Pattern Recognition 113, p. 107760. Cited by: §1.