2024
Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment
ICLR 2024poster
We introduce a method to train vision-language models for remote-sensing images without using any textual annotations. Our key insight is to use co-located internet imagery taken on the ground as an intermediary for connecting remote-sensing images and language. Specifically, we train an image enco…