Skip to main navigation Skip to search Skip to main content

GeoPriorclip: a foundational remote sensing vision-language model enhanced with cascaded geographic information priors

  • Aokun Liang
  • , Xi Xiao
  • , Xiangyun Hu
  • , Tao Ke
  • , Tianyang Wang
  • , Yibing Xiong
  • , Huiwei Jiang
  • , Xiao Wang

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

Remote sensing vision–language models (RSVLMs) have made notable progress in bridging the semantic gap between satellite imagery and natural language. However, two fundamental limitations persist. First, existing vision–language corpora for remote sensing are constructed using general-purpose models and lack integration of structured geographic priors from authoritative remote sensing resources. Second, current RSVLMs do not explicitly model geometric boundaries or spatial relations, leading to suboptimal image–text alignment. To address these limitations, we introduce GeoPrior, a large-scale tri-modal dataset comprising 828k satellite images, 2.48 M textual descriptions, and aligned rasterized maps. GeoPrior encodes detailed geographic priors–including geometric structures, topological relationships, and semantic attributes–extracted from authoritative vector maps. These priors guide GPT-4o in generating knowledge-rich captions tailored to remote sensing, while the rasterized maps provide an additional modality that captures fine-grained boundary information beyond the expressiveness of natural language. Building upon GeoPrior, we propose GeoPriorCLIP, a vision–language model tailored for remote sensing. Our key technical contribution is the geo-aware cross-modal attention module, which injects map-derived spatial priors into the CLIP image encoder to enhance visual representation with explicit geometric and topological awareness. Extensive experiments on 18 public benchmark datasets across four tasks–zero-shot classification, cross-modal retrieval, semantic localization, and zero-shot semantic segmentation–demonstrate that GeoPriorCLIP consistently outperforms state-of-the-art RSVLMs. The code and data will be released upon acceptance.

Original languageEnglish
JournalGeo-Spatial Information Science
DOIs
StateAccepted/In press - 2026

Keywords

  • Remote Sensing
  • Vision-language model
  • contrastive learning
  • multimodal learning

Fingerprint

Dive into the research topics of 'GeoPriorclip: a foundational remote sensing vision-language model enhanced with cascaded geographic information priors'. Together they form a unique fingerprint.

Cite this