GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data
Abstract
Precise spatial understanding in Earth Observation is essen-tial for translating raw aerial imagery into actionable insights for criti-cal applications like urban planning, environmental monitoring and dis-aster management. However, Multimodal Large Language Models ex-hibit critical deficiencies in fine-grained spatial understanding withinRemote Sensing, primarily due to a reliance on limited or repurposedlegacy datasets. To bridge this gap, we introduce a large-scale datasetgrounded in verifiable cadastral vector data, comprising 3.8 million an-notated objects across 510k high-resolution images with 135 granularsemantic categories. We validate this resource through a comprehensiveinstruction-tuning benchmark spanning seven spatial grounding tasks.Our evaluation establishes a robust baseline using a standard LLaVAarchitecture. We show that while current RS-specialized and commer-cial models (e.g., Gemini) struggle in zero-shot settings, high-fidelity su-pervision effectively bridges this gap, enabling standard architecturesto master fine-grained spatial grounding without complex architecturalmodifications. Data, pretrained model and code are available at: https://huggingface.co/datasets/RogerFerrod/GroundSet