ICASSP 2025accepted0 citations

Fusing Multimodality of Large Language Models and Satellite Imagery via Simplicial Contrastive Learning for Latent Urban Feature Identification and Environmental Application

Yuzhou Chen, Jiue-An Yang, Hugo Kyo Lee, Calvin P. Tribby, Tarik Benmarhnia, Marta M. Jankowska, Yulia R. Gel

Abstract

Satellite imagery is a readily available data source for monitoring a broad range of urban geographical contexts related to environmental, socio-demographic, and health disparities. To analyze satellite images, deep learning (DL) tools efficiently extract latent multi-dimensional characteristics, beyond identifying specific urban elements like roads and houses. However, current DL approaches tend to largely rely on Convolutional Neural Networks applied to high-resolution imagery, and as such may be limited to capturing only local contextual information. To address this fundamental limitation, we propose to fuse the modalities of satellite imagery and a large language model (LLM). In particular, we develop a novel LLM-based Simplicial Contrastive Learning model (LLM-SCL) based on mutual information maximization between the latent simplicial complex-level representations of two kinds of augmented (superpixel) graphs, which allows for cohesive integration of LLM prompts and learning of both local and global higher-order properties of satellite imagery (from all pixels in an image). Extensive experiments on satellite imagery at several resolutions in Tijuana, Mexico, Los Angeles and San Diego, USA, suggest that LLM-SCL significantly outperforms state-of-the-art baselines on unsupervised image classification tasks. As such, the proposed LLM-SCL opens a new path for more accurate evaluations of latent urban forms and their associations with environmental and health outcome disparities.

BibTeX
@inproceedings{icassp2025_fusingmultimodal,
  title = {Fusing Multimodality of Large Language Models and Satellite Imagery via Simplicial Contrastive Learning for Latent Urban Feature Identification and Environmental Application},
  author = {Yuzhou Chen and Jiue-An Yang and Hugo Kyo Lee and Calvin P. Tribby and Tarik Benmarhnia and Marta M. Jankowska and Yulia R. Gel},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Fusing Multimodality of Large Language Models and Satellite Imagery via Simplicial Contrastive Learning for Latent Urban Feature Identification and Environmental Application · ICASSP 2025