← Search

Linyi Jin

11 accepted papers

2026

Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation

ICLR 2026poster

Camera-centric understanding and generation are two cornerstones of spatial intelligence, yet they are typically studied in isolation. We present Puffin, a unified camera-centric multimodal model that extends spatial awareness along the camera dimension. Puffin integrates language regression and dif…

Cited by 0SourcecodeScholar
2025

MegaSaM: Accurate, Fast and Robust Structure and Motion from Casual Dynamic Videos

CVPR 2025award

We present a system that allows for accurate, fast, and robust estimation of camera parameters and depth maps from casual monocular videos of dynamic scenes. Most conventional structure from motion and monocular SLAM techniques assume input videos that feature predominantly static scenes with large…

Cited by 18SourcePDFScholar
2025

Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos

CVPR 2025poster

Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly supervising methods for recovering 3D motion remains challenging due…

2024

FAR: Flexible Accurate and Robust 6DoF Relative Camera Pose Estimation

CVPR 2024highlight

Estimating relative camera poses between images has been a central problem in computer vision. Methods that find correspondences and solve for the fundamental matrix offer high precision in most cases. Conversely methods predicting pose directly using neural networks are more robust to limited overl…

Cited by 4SourcePDFScholar
2023

Learning To Predict Scene-Level Implicit 3D From Posed RGBD Data

CVPR 2023poster

We introduce a method that can learn to predict scene-level implicit functions for 3D reconstruction from posed RGBD data. At test time, our system maps a previously unseen RGB image to a 3D reconstruction of a scene via implicit functions. While implicit functions for 3D reconstruction have often b…

Cited by 2SourcePDFScholar
2023

Perspective Fields for Single Image Camera Calibration

CVPR 2023highlight

Geometric camera calibration is often required for applications that understand the perspective of the image. We propose perspective fields as a representation that models the local perspective properties of an image. Perspective Fields contain per-pixel information about the camera view, parameteri…

2022

PlaneFormers: From Sparse View Planes to 3D Reconstruction

ECCV 2022poster

"We present an approach for the planar surface reconstruction of a scene from images with limited overlap. This reconstruction task is challenging since it requires jointly reasoning about single image 3D reconstruction, correspondence between images, and the relative camera pose between images. Pas…

2022

Understanding 3D Object Articulation in Internet Videos

CVPR 2022poster

We propose to investigate detecting and characterizing the 3D planar articulation of objects from ordinary RGB videos. While seemingly easy for humans, this problem poses many challenges for computers. Our approach is based on a top-down detection system that finds planes that can be articulated. Th…

Cited by 24PDFcodeScholar