← Search

Li Mi

9 accepted papers

2026

GeoFAR: Geography-Informed Frequency-Aware Super-Resolution for Climate Data

ICLR 2026poster

Super-resolving climate data is crucial for fine-grained decision-making in various domains, ranging from agriculture to environmental conservation. However, existing super-resolution approaches struggle to generate the high-frequency spatial information present in climate data, especially over regi…

Cited by 0SourceScholar
2026

SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs

ICML 2026poster

Despite recent successes, *test-time scaling* $-$i.e., dynamically expanding the token budget during inference as needed$-$ remains brittle for vision-language models (VLMs): unstructured chains-of-thought about images entangle perception and reasoning, leading to long, disorganized contexts where s…

Cited by 0SourceScholar
2026

Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

ICML 2026poster

Multimodal LLMs lack a systematic understanding of visual dynamics in complex human world activities, which requires the model to predict or simulate multiple levels of dynamic constituents, such as the general progression of actions and the associated changes of low-level details in the world. To a…

Cited by 0SourceScholar
2025

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration

ICCV 2025poster

Active Geo-localization (AGL) is the task of localizing a goal, represented in various modalities (e.g., aerial images, ground-level images, or text), within a predefined search area. Current methods approach AGL as a goal-reaching reinforcement learning (RL) problem with a distance-based reward. Th…

Cited by 0SourcePDFScholar
2025

VinaBench: Benchmark for Faithful and Consistent Visual Narratives

CVPR 2025poster

Visual narrative generation transforms textual narratives into sequences of images illustrating the content of the text. However, generating visual narratives that are faithful to the input text and self-consistent across generated images remains an open challenge, due to the lack of knowledge const…

Cited by 1SourcePDFScholar
2024

ConGeo: Robust Cross-view Geo-localization across Ground View Variations

ECCV 2024poster

"∗ Equal contribution Corresponding author (xuchangeis@whu.edu.cn) Cross-view geo-localization aims at localizing a ground-level query image by matching it to its corresponding geo-referenced aerial view. In real-world scenarios, the task requires accommodating diverse ground images captured by user…

Cited by 7SourcePDFScholar
2024

ConVQG: Contrastive Visual Question Generation with Multimodal Guidance

AAAI 2024technical

Asking questions about visual environments is a crucial way for intelligent agents to understand rich multi-faceted scenes, raising the importance of Visual Question Generation (VQG) systems. Apart from being grounded to the image, existing VQG systems can use textual constraints, such as expected a…