← Search

Zitian Tang

5 accepted papers

2026

Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals

CVPR 2026

Recent advancements in video generation have enabled the development of "world models" capable of simulating potential futures for robotics and planning. However, specifying precise goals for these models remains a challenge; text instructions are often too abstract to capture physical nuances, whil

Cited by 0SourcecodeScholar
2026

MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-Correction

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently demonstrated promising capabilities in multimodal coding tasks such as chart-to-code generation. However, existing methods primarily rely on supervised fine-tuning (SFT), which requires the model to learn code patterns through chart-code pairs b

Cited by 0SourceScholar
2025

How Can Objects Help Video-Language Understanding?

ICCV 2025poster

Do we still need to represent objects explicitly in multimodal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which objects and spatiotemporal relationships may be implicitly modeled. To the other extreme, image captions by themselves prov…

2025

SAGE: A Unified Framework for Generalizable Object State Recognition with State-Action Graph Embedding

NeurIPS 2025oral

Recognizing the physical states of objects and their transformations within videos is crucial for structured video understanding and enabling robust real-world applications, such as robotic manipulation. However, pretrained vision-language models often struggle to capture these nuanced dynamics and…

Cited by 0SourceScholar
2023

What Happened 3 Seconds Ago? Inferring the Past With Thermal Imaging

CVPR 2023poster

Inferring past human motion from RGB images is challenging due to the inherent uncertainty of the prediction problem. Thermal images, on the other hand, encode traces of past human-object interactions left in the environment via thermal radiation measurement. Based on this observation, we collect th…