The advent of Geospatial Foundation Models (GeoFMs) marks a new era in the analysis of satellite and aerial imagery. This paper describes GeoFMs as AI/ML models pre-trained on massive geospatial datasets through various methodologies.
The core paradigm shift introduced by GeoFMs is the separation of duties, allowing large-scale model providers to perform computationally intensive pretraining while domain experts fine-tune or prompt these models for specific, mission-critical tasks. This democratizes access to cutting-edge AI/ML while maintaining the security and confidentiality of downstream tasks.
We explore the novel capabilities unlocked by different types of GeoFMs, distinguishing between finetunable vision models produced by self-supervised techniques like masked auto-encoding and vision-language models produced by contrastive learning, which enable zero-shot tasks such as open-vocabulary image analysis.
We then discuss practical considerations for operationalizing GeoFMs, from performance-cost analysis to the broader MLOps ecosystem. A taxonomy of model adaptation strategies is introduced, proposing a framework for domain experts to select the most cost-effective adaptation approach for their specific missions.
Finally, we present a forward-looking vision of Agentic Geospatial Reasoning, where Large Language Models act as intelligent orchestrators, leveraging GeoFMs as tools to answer high-level user queries in natural language and automate complex analytical workflows, shifting the field from perception to cognition.