← Search

Zongzhao Li

5 accepted papers

2026

Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task Progress

CVPR 2026

Accurately estimating task progress is critical for embodied agents to plan and execute long-horizon, multi-step tasks. Despite promising advances, existing Vision-Language Models (VLMs) based methods primarily leverage their video understanding capabilities, while neglecting their complex reasoning

Cited by 0SourceScholar
2026

STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs

CVPR 2026

Multimodal Large Language Models (MLLMs) remain far from human-level performance in multi-view spatial reasoning, where models must establish object correspondences across view and infer coherent scene semantics. We analyze this limitation through the Transformation-Driven Visual Reasoning (TVR) tas

Cited by 0SourcecodeScholar
2025

Large Language-Geometry Model: When LLM meets Equivariance

ICML 2025poster

Accurately predicting 3D structures and dynamics of physical systems is crucial in scientific applications. Existing approaches that rely on geometric Graph Neural Networks (GNNs) effectively enforce $\mathrm{E}(3)$-equivariance, but they often fail in leveraging extensive broader information. While…

Cited by 4SourcePDFScholar
2025

Size-Generalizable RNA Structure Evaluation by Exploring Hierarchical Geometries

ICLR 2025poster

Understanding the 3D structure of RNA is essential for deciphering its function and developing RNA-based therapeutics. Geometric Graph Neural Networks (GeoGNNs) that conform to the $\mathrm{E}(3)$-symmetry have advanced RNA structure evaluation, a crucial step toward RNA structure prediction. Howeve…

Cited by 2SourcePDFScholar
2022

Deconfounding Physical Dynamics with Global Causal Relation and Confounder Transmission for Counterfactual Prediction

AAAI 2022technical

Discovering the underneath causal relations is the fundamental ability for reasoning about the surrounding environment and predicting the future states in the physical world. Counterfactual prediction from visual input, which requires simulating future states based on unrealized situations in the pa…

Cited by 5SourcePDFScholar