Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos
Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, which are unable to generalize to complex real-world settin