OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization
Abstract
Fine-grained Cross-View Localization (CVL) estimates theprecise position and orientation of a ground-level image by aligningit with geo-referenced aerial imagery, offering a scalable alternative toGlobal Navigation Satellite Systems (GNSS) in challenging urban envi-ronments. Existing datasets rely on data collected with high-end sensorsuites, which inherently limit image diversity and scalability. While in-the-wild images are abundant, their noisy geo-tags make them unsuit-able for reliable evaluation. To bridge this gap, we introduce OpenCVL,a large-scale, diverse, and open dataset containing 617,388 ground-aerialimage pairs spanning 41 cities across four European countries. All imagesare sourced from permissive platforms, ensuring long-term accessibilityand supporting open and reproducible research. The training set com-bines images captured with high-end sensors with diverse in-the-wild im-agery. We further develop a data curation framework that filters and cor-rects pose annotations to construct reliable in-the-wild evaluation data.In addition, OpenCVL includes dedicated cross-area and snowy test setsto assess generalization and robustness. Experiments with a state-of-the-art CVL model on OpenCVL show that incorporating noisy in-the-wilddata consistently improves performance on clean test sets, suggesting apromising direction for scaling CVL with diverse real-world imagery.