← Search

Zhenjie Mao

3 accepted papers

2026

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning

ICML 2026poster

Spatial reasoning from egocentric videos is inherently challenging because the observable evidence is constrained by the camera trajectory. Existing methods perform spatial reasoning in a single inference pass, forcing models to resolve geometric ambiguity through semantic priors rather than verifia…

Cited by 0SourceScholar
2025

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition

ICML 2025poster

Video understanding is a complex challenge that requires effective modeling of spatial-temporal dynamics. With the success of image foundation models (IFMs) in image understanding, recent approaches have explored parameter-efficient fine-tuning (PEFT) to adapt IFMs for video. However, most of the…

Cited by 0SourcePDFScholar
2025

SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation

NeurIPS 2025poster

Referring Image Segmentation (RIS) aims to segment the target object in an image given a natural language expression. While recent methods leverage pre-trained vision backbones and more training corpus to achieve impressive results, they predominantly focus on simple expressions—short, clear noun ph…

Cited by 0SourceScholar