OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
Henry Herzog, Favyen Bastani, Yawen Zhang, Gabriel Tseng, Joseph Redmon, Hadrien Sablon, Ryan Park, Jacob Morrison
Abstract
Earth observation data presents a unique challenge: it is spatial like images, sequential like video or text, and highly multimodal. We present Helios: a multimodal, spatio-temporal foundation model that employs a novel self-supervised learning formulation, masking strategy, and loss all designed for the Earth observation domain. Helios achieves state-of-the-art performance compared to 12 other foundation models across a variety of research benchmarks and real-world tasks from external partners. When evaluating embeddings Helios achieves the best performance on 15 out of 24 tasks, and with full fine-tuning it is the best on 19 of 29 tasks. We deploy Helios as the backbone of an end-to-end platform for data collection, labeling, training, and inference of Earth observation models. The Helios platform puts frontier foundation models and powerful data management tools into the hands of non-profits and NGOs working to solve the world's biggest problems. Helios source code, training data, and pre-trained weights are available at REDACTED.
BibTeX
@inproceedings{cvpr2026_olmoearthstablel,
title = {OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation},
author = {Henry Herzog and Favyen Bastani and Yawen Zhang and Gabriel Tseng and Joseph Redmon and Hadrien Sablon and Ryan Park and Jacob Morrison and Alexandra Buraczynski and Karen Farley and Josh Hansen and Andrew Howe and Patrick Alan Johnson and Mark Otterlee and Ted Schmitt and Hunter Pitelka and Stephen Daspit and Rachel Ratner and Christopher Wilhelm and Sebastian Wood and Mike Jacobi and Hannah Kerner and Evan Shelhamer and Ali Farhadi and Ranjay Krishna and Patrick Beukema},
booktitle = {CVPR 2026},
year = {2026}
}