← Search

Jason Y. Zhang

14 accepted papers

2026

VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression

CVPR 2026

Evaluations of image compression performance which include human preferences have generally found that naive distortion functions such as MSE are insufficiently aligned to human perception.In order to align compression models to human perception, prior work has employed differentiable perceptual los

Cited by 0SourceScholar
2025

Bolt3D: Generating 3D Scenes in Seconds

ICCV 2025poster

We present a latent diffusion model for fast feed-forward 3D scene generation. Given one or more images, our model Bolt3D directly samples a 3D scene representation in less than seven seconds on a single GPU. We achieve this by leveraging powerful and scalable existing 2D diffusion network architect…

2025

Can Generative Video Models Help Pose Estimation?

CVPR 2025highlight

Pairwise pose estimation from images with little or no overlap is an open challenge in computer vision. Existing methods, even those trained on large-scale datasets, struggle in these scenarios due to the lack of identifiable correspondences or visual overlap. Inspired by the human ability to infer…

2025

DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion

CVPR 2025poster

Current Structure-from-Motion (SfM) methods typically follow a two-stage pipeline, combining learned or geometric pairwise reasoning with a subsequent global optimization step. In contrast, we propose a data-driven multi-view reasoning approach that directly infers 3D scene geometry and camera poses…

2025

UnCommon Objects in 3D

CVPR 2025poster

We introduce Uncommon Objects in 3D (uCO3D), a new object-centric dataset for 3D deep learning and 3D generative AI. uCO3D is the largest publicly-available collection of high-resolution videos of objects with 3D annotations that ensures full-360 degree coverage. uCO3D is significantly more diverse…

2024

Cameras as Rays: Pose Estimation via Ray Diffusion

ICLR 2024oral

Estimating camera poses is a fundamental task for 3D reconstruction and remains challenging given sparsely sampled views (<10). In contrast to existing approaches that pursue top-down prediction of global parametrizations of camera extrinsics, we propose a distributed representation of camera pose t…

Cited by 60SourcePDFScholar
2023

SparsePose: Sparse-View Camera Pose Regression and Refinement

CVPR 2023poster

Camera pose estimation is a key step in standard 3D reconstruction pipelines that operates on a dense set of images of a single object or scene. However, methods for pose estimation often fail when there are only a few images available because they rely on the ability to robustly identify and match…

Cited by 45SourcePDFScholar
2022

RelPose: Predicting Probabilistic Relative Rotation for Single Objects in the Wild

ECCV 2022poster

"We describe a data-driven method for inferring the camera viewpoints given multiple images of an arbitrary object. This task is a core component of classic geometric pipelines such as SfM and SLAM, and also serves as a vital pre-processing requirement for contemporary neural approaches (e.g. NeRF)…

Cited by 94SourcePDFScholar
2021

NeRS: Neural Reflectance Surfaces for Sparse-view 3D Reconstruction in the Wild

NeurIPS 2021poster

Recent history has seen a tremendous growth of work exploring implicit representations of geometry and radiance, popularized through Neural Radiance Fields (NeRF). Such works are fundamentally based on a (implicit) {\em volumetric} representation of occupancy, allowing them to model diverse scene s…

2020

Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild

ECCV 2020poster

We present a method that infers spatial arrangements and shapes of humans and objects in a globally consistent 3D scene, all from a single image in-the-wild captured in an uncontrolled environment. Notably, our method runs on datasets without any scene- or object-level 3D supervision. Our key insigh…