← Search

Patrick Rim

10 accepted papers

2026

Iris: Integrating Language into Diffusion-based Monocular Depth Estimation

CVPR 2026

Conventional monocular depth estimators suffer from visual ambiguities and nuisances. We demonstrate that language can improve the fidelity of estimates by providing additional information through text as a condition, thereby reducing the solution space for depth estimates. This conditional distribu

Cited by 0SourceScholar
2026

ODE-GS: Latent ODEs for Dynamic Scene Extrapolation with 3D Gaussian Splatting

ICLR 2026poster

We introduce ODE-GS, a novel approach that integrates 3D Gaussian Splatting with latent neural ordinary differential equations (ODEs) to enable future extrapolation of dynamic 3D scenes. Unlike existing dynamic scene reconstruction methods, which rely on time-conditioned deformation networks and are…

Cited by 0SourcecodeScholar
2026

ORCaS: Unsupervised Depth Completion via Occluded Region Completion as Supervision

ICLR 2026poster

We propose a method for inferring an egocentric dense depth map from an RGB image and a sparse point cloud. The crux of our method lies in modeling the 3D scene implicitly within the latent space and learning an inductive bias in an unsupervised manner through principles of Structure-from-Motion. T…

Cited by 0SourceScholar
2026

SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild

CVPR 2026

Accurate 3D understanding of human hands and objects during manipulation remains a significant challenge for egocentric computer vision. Existing hand-object interaction datasets are predominantly captured in controlled studio settings, which limits both environmental diversity and the ability of mo

Cited by 0SourcecodeScholar
2025

ETA: Energy-based Test-time Adaptation for Depth Completion

ICCV 2025poster

We propose a method of adapting pretrained depth completion models to test time data in an unsupervised manner. Depth completion models are (pre)trained to produce dense depth maps from pairs of RGB image and sparse depth maps in ideal capture conditions (source domain), e.g., well-illuminated, high…

Cited by 0SourcePDFScholar
2025

Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens

ICCV 2025accepted

We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images. Despite being trained on tens of millions of images, FMDEs are susceptible to the covariate shift introduced by changes in camera calibration (intrinsic, distortion) param…

2025

ProtoDepth: Unsupervised Continual Depth Completion with Prototypes

CVPR 2025poster

We present ProtoDepth, a novel prototype-based approach for continual learning of unsupervised depth completion, the multimodal 3D reconstruction task of predicting dense depth maps from RGB images and sparse point clouds. The unsupervised learning paradigm is well-suited for continual learning, as…

Cited by 1SourcePDFScholar
2023

Quadric Representations for LiDAR Odometry, Mapping and Localization

RA-L 2023

Current LiDAR odometry, mapping and localization methods leverage point-wise representations of 3D scenes and achieve high accuracy in autonomous driving tasks. However, the space-inefficiency of methods that use point-wise representations limits their development and usage in practical applications

Cited by 13SourceScholar
2023

SparseFusion: Fusing Multi-Modal Sparse Representations for Multi-Sensor 3D Object Detection

ICCV 2023poster

By identifying four important components of existing LiDAR-camera 3D object detection methods (LiDAR and camera candidates, transformation, and fusion outputs), we observe that all existing methods either find dense candidates or yield dense representations of scenes. However, given that objects occ…

Cited by 77PDFcodeScholar