← Search

Manqi Zhao

3 accepted papers

2026

Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Model

RSS 2026poster

Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for evaluating Vision-Language-Action (VLA) systems. However, their reliability as evaluation benchmarks is often limited by the representation gap between low-dimensional actions and high-dimensional video synthesis. Thi…

Cited by 0SourceScholar
2026

ReAttnCLIP: Training-Free Open-Vocabulary Remote Sensing Image Segmentation via Re-defined Attention in CLIP

CVPR 2026

Remote sensing image segmentation is critical for a range of applications, including natural disaster monitoring and precision agriculture. Open-vocabulary segmentation enhances flexibility by removing fixed category constraints, enabling more fine-grained and adaptive scene understanding. Unlike CL

Cited by 0SourceScholar
2026

SAM2MOT: A Novel Paradigm of Multi-Object Tracking by Segmentation

AAAI 2026technical

Inspired by Segment Anything 2, which generalizes segmentation from images to videos, we propose SAM2MOT—a novel segmentation-driven paradigm for multi-object tracking that breaks away from the conventional detection-association framework. In contrast to previous approaches that treat segmentation a

Cited by 0SourcePDFScholar