← Search

Utkarsh Mall

13 accepted papers

2026

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

CVPR 2026

The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus predominantly on coarse-grained tasks such as action segmentation, classification, captioning, and retrieval. Furthermore, these benchmarks often rely

Cited by 0SourcecodeScholar
2025

DiSciPLE: Learning Interpretable Programs for Scientific Visual Discovery

CVPR 2025poster

Visual data is used in numerous different scientific workflows ranging from remote sensing to ecology. As the amount of observation data increases, the challenge is not just to make accurate predictions but also to understand the underlying mechanisms for those predictions. Good interpretation is im…

Cited by 0SourcePDFScholar
2025

MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing

NeurIPS 2025spotlight

Natural disasters cause devastating damage to communities and infrastructure every year. Effective disaster response is hampered by the difficulty of accessing affected areas during and after events. Remote sensing has allowed us to monitor natural disasters in a remote way. More recently there have…

Cited by 0SourceScholar
2025

Scale-aware Recognition in Satellite Images under Resource Constraints

ICLR 2025poster

Recognition of features in satellite imagery (forests, swimming pools, etc.) depends strongly on the spatial scale of the concept and therefore the resolution of the images. This poses two challenges: Which resolution is best suited for recognizing a given concept, and where and when should the cost…

Cited by 0SourcePDFScholar
2024

AllClear: A Comprehensive Dataset and Benchmark for Cloud Removal in Satellite Imagery

NeurIPS 2024poster

Clouds in satellite imagery pose a significant challenge for downstream applications. A major challenge in current cloud removal research is the absence of a comprehensive benchmark and a sufficiently large and diverse training dataset. To address this problem, we introduce the largest public datase…

2024

Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment

ICLR 2024poster

We introduce a method to train vision-language models for remote-sensing images without using any textual annotations. Our key insight is to use co-located internet imagery taken on the ground as an intermediary for connecting remote-sensing images and language. Specifically, we train an image enco…

Cited by 50SourcePDFScholar
2023

Change-Aware Sampling and Contrastive Learning for Satellite Images

CVPR 2023poster

Automatic remote sensing tools can help inform many large-scale challenges such as disaster management, climate change, etc. While a vast amount of spatio-temporal satellite image data is readily available, most of it remains unlabelled. Without labels, this data is not very useful for supervised le…

2022

Change Event Dataset for Discovery from Spatio-temporal Remote Sensing Imagery

NeurIPS 2022accept

Satellite imagery is increasingly available, high resolution, and temporally detailed. Changes in spatio-temporal datasets such as satellite images are particularly interesting as they reveal the many events and forces that shape our world. However, finding such interesting and meaningful change e…

Cited by 17SourcePDFScholar
2021

PiCIE: Unsupervised Semantic Segmentation Using Invariance and Equivariance in Clustering

CVPR 2021poster

We present a new framework for semantic segmentation without annotations via clustering. Off-the-shelf clustering methods are limited to curated, single-label, and object-centric images yet real-world data are dominantly uncurated, multi-label, and scene-centric. We extend clustering from images to…

Cited by 238PDFcodeScholar