← Search

Wanhee Lee

5 accepted papers

2026

Perceptual 3D Simulation With Physical World Modeling

CVPR 2026

Predicting how a scene will evolve after a desired 3D transformation from images is a central goal in vision, graphics, and robotics. Yet unlike ideal simulators with full access to 3D geometry and dynamics, real world systems must rely on perceptual inputs and local actions that are inherently part

Cited by 0SourceScholar
2026

Physical Object Understanding with a Physically Controllable World Model

CVPR 2026

A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that govern their interactions. Solving these tasks requires world models capable of inferring distributional states of the world from partial observations -

Cited by 0SourceScholar
2026

Unified 3D Scene Understanding Through Physical World Modeling

ICLR 2026poster

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essential for perception and interaction. Existing approaches have typically addressed these tasks in isolation, preventing th…

Cited by 0SourceScholar
2025

Taming generative video models for zero-shot optical flow extraction

NeurIPS 2025poster

Extracting optical flow from videos remains a core computer vision problem. Motivated by the recent success of large general-purpose models, we ask whether frozen self-supervised video models trained only to predict future frames can be prompted, without fine-tuning, to output flow. Prior attempts t…

Cited by 0SourceScholar
2024

Understanding Physical Dynamics with Counterfactual World Modeling

ECCV 2024poster

"The ability to understand physical dynamics is critical for agents to act in the world. Here, we use Counterfactual World Modeling (CWM) to extract vision structures for dynamics understanding. CWM uses a temporally-factored masking policy for masked prediction of video data without annotations. Th…