← Search

Siyuan Zhou

19 accepted papers

2026

OPTION: Optimal Transport–Guided Flow Matching for Incomplete and Unaligned Multi-View Clustering

ICML 2026poster

Multi-view clustering effectively exploits rich information from multiple views, yet real-world applications are frequently challenged by missing views and cross-view sample misalignment, hindering cross-view modeling and resulting inferior clustering performance. To address these challenges, this p…

Cited by 0SourceScholar
2026

PhysInOne: Visual Physics Learning and Reasoning in One Suite

CVPR 2026

We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike existing datasets limited to merely hundreds or thousands of examples, PhysInOne provides 2 million videos across 153,810 dynamic 3D scenes, covering 71

Cited by 0SourcecodeScholar
2025

AdaWorld: Learning Adaptable World Models with Latent Actions

ICML 2025poster

World models aim to learn action-controlled future prediction and have proven essential for the development of intelligent agents. However, most existing world models rely heavily on substantial action-labeled data and costly training, making it challenging to adapt to novel environments with hetero…

2025

FreeGave: 3D Physics Learning from Dynamic Videos by Gaussian Velocity

CVPR 2025poster

In this paper, we aim to model 3D scene geometry, appearance, and the underlying physics purely from multi-view videos. By applying various governing PDEs as PINN losses or incorporating physics simulation into neural networks, existing works often fail to learn complex physical motions at boundarie…

2025

Learning 3D Persistent Embodied World Models

NeurIPS 2025poster

The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing work has explored how to construct such world models using vid…

Cited by 0SourceScholar
2025

Learning 4D Embodied World Models

ICCV 2025poster

This paper presents an effective approach for learning novel 4D embodied world models, which predict the dynamic evolution of 3D scenes over time in response to an embodied agent's actions, providing both spatial and temporal consistency. We propose to learn a 4D world model by training on RGB-DN (R…

2025

MindJourney: Test-Time Scaling with World Models for Spatial Reasoning

NeurIPS 2025poster

Spatial reasoning in 3D space is central to human cognition and indispensable for embodied tasks such as navigation and manipulation. However, state-of-the-art vision–language models (VLMs) struggle frequently with tasks as simple as anticipating how a scene will look after an egocentric motion: the…

Cited by 0SourceScholar
2025

RayletDF: Raylet Distance Fields for Generalizable 3D Surface Reconstruction from Point Clouds or Gaussians

ICCV 2025poster

In this paper, we present a generalizable method for 3D surface reconstruction from raw point clouds or pre-estimated 3D Gaussians by 3DGS from RGB images. Unlike existing coordinate-based methods which are often computationally intensive when rendering explicit surfaces, our proposed method, named…

Cited by 0SourcePDFScholar
2024

RoboDreamer: Learning Compositional World Models for Robot Imagination

ICML 2024poster

Text-to-video models have demonstrated substantial potential in robotic decision-making, enabling the imagination of realistic plans of future actions as well as accurate environment simulation. However, one major issue in such models is generalization -- models are limited to synthesizing videos su…

Cited by 23SourcePDFScholar
2023

Adaptive Online Replanning with Diffusion Models

NeurIPS 2023poster

Diffusion models have risen a promising approach to data-driven planning, and have demonstrated impressive robotic control, reinforcement learning, and video planning performance. Given an effective planner, an important question to consider is replanning -- when given plans should be regenerated du…

Cited by 22SourcePDFScholar
2022

Finding Fallen Objects via Asynchronous Audio-Visual Integration

CVPR 2022poster

The way an object looks and sounds provide complementary reflections of its physical properties. In many settings cues from vision and audition arrive asynchronously but must be integrated, as when we hear an object dropped on the floor and then must find it. In this paper, we introduce a setting in…

Cited by 20PDFScholar
2022

The ThreeDWorld Transport Challenge: A Visually Guided Task-and-Motion Planning Benchmark Towards Physically Realistic Embodied AI

ICRA 2022poster

We introduce a visually-guided task-and-motion planning benchmark, which we call the ThreeDWorld Trans-port Challenge. In this challenge, an embodied agent is spawned randomly in a simulated physical home environment and required to transport a small set of objects scattered around the house with co…

Cited by 46SourceScholar
2022

Weak-shot Semantic Segmentation via Dual Similarity Transfer

NeurIPS 2022accept

Semantic segmentation is a practical and active task, but severely suffers from the expensive cost of pixel-level labels when extending to more classes in wider applications. To this end, we focus on the problem named weak-shot semantic segmentation, where the novel classes are learnt from cheaper i…

2021

Learning Task Decomposition with Ordered Memory Policy Network

ICLR 2021poster

Many complex real-world tasks are composed of several levels of subtasks. Humans leverage these hierarchical structures to accelerate the learning process and achieve better generalization. In this work, we study the inductive bias and propose Ordered Memory Policy Network (OMPN) to discover subtask…

Cited by 21SourcePDFScholar
2021

Native Chinese Reader: A Dataset Towards Native-Level Chinese Machine Reading Comprehension

NeurIPS 2021poster

We present Native Chinese Reader (NCR), a new machine reading comprehension MRC) dataset with particularly long articles in both modern and classical Chinese. NCR is collected from the exam questions for the Chinese course in China’s high schools, which are designed to evaluate the language profic…

Cited by 2SourceScholar
2021

PlasticineLab: A Soft-Body Manipulation Benchmark with Differentiable Physics

ICLR 2021spotlight

Simulated virtual environments serve as one of the main driving forces behind developing and evaluating skill learning algorithms. However, existing environments typically only simulate rigid body physics. Additionally, the simulation process usually does not provide gradients that might be useful f…

2019

Transferable Interactiveness Knowledge for Human-Object Interaction Detection

CVPR 2019poster

Human-Object Interaction (HOI) Detection is an important problem to understand how humans interact with objects. In this paper, we explore Interactiveness Knowledge which indicates whether human and object interact with each other or not. We found that interactiveness knowledge can be learned across…

Cited by 383PDFcodeScholar