← Search

Khai Loong Aw

4 accepted papers

2026

Unified 3D Scene Understanding Through Physical World Modeling

ICLR 2026poster

Understanding 3D scenes requires flexible combinations of visual reasoning tasks, including depth estimation, novel view synthesis, and object manipulation, all of which are essential for perception and interaction. Existing approaches have typically addressed these tasks in isolation, preventing th…

Cited by 0SourceScholar
2025

Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals

NeurIPS 2025spotlight

Estimating motion primitives from video (e.g., optical flow and occlusion) is a critically important computer vision problem with many downstream applications, including controllable video generation and robotics. Current solutions are primarily supervised on synthetic data or require tuning of situ…

Cited by 0SourceScholar
2025

Taming generative video models for zero-shot optical flow extraction

NeurIPS 2025poster

Extracting optical flow from videos remains a core computer vision problem. Motivated by the recent success of large general-purpose models, we ask whether frozen self-supervised video models trained only to predict future frames can be prompted, without fine-tuning, to output flow. Prior attempts t…

Cited by 0SourceScholar