← Search

Zhoutong Zhang

23 accepted papers

2026

Generative Video Motion Editing with 3D Point Tracks

CVPR 2026

Camera and object motions are central to a video's narrative. However, precisely editing these captured motions remains a significant challenge, especially under complex object movements. Current motion-controlled image-to-video (I2V) approaches often lack full-scene context for consistent video edi

Cited by 0SourceScholar
2026

Hist2Style: Histogram-Guided Stylization with Bilateral Grids

CVPR 2026

Photorealistic style transfer aims to match the color and tone of an input image to that of a style target while preserving the content and details of the original scene. Although existing large image models can facilitate these kinds of appearance edits, their high computational demands, potential

Cited by 0SourcecodeScholar
2025

Classic Video Denoising in a Machine Learning World: Robust, Fast, and Controllable

CVPR 2025poster

Denoising is a crucial step in many video processing pipelines such as in interactive editing, where high quality, speed, and user control are essential. While recent approaches achieve significant improvements in denoising quality by leveraging deep learning, they are prone to unexpected failures d…

Cited by 0SourcePDFScholar
2024

DriveTrack: A Benchmark for Long-Range Point Tracking in Real-World Videos

CVPR 2024poster

This paper presents DriveTrack a new benchmark and data generation framework for long-range keypoint tracking in real-world videos. DriveTrack is motivated by the observation that the accuracy of state-of-the-art trackers depends strongly on visual attributes around the selected keypoints such as te…

Cited by 9SourcePDFScholar
2024

Fast View Synthesis of Casual Videos with Soup-of-Planes

ECCV 2024poster

"Novel view synthesis from an in-the-wild video is difficult due to challenges like scene dynamics and lack of parallax. While existing methods have shown promising results with implicit neural radiance fields, they are slow to train and render. This paper revisits explicit video representations to…

2024

FeatUp: A Model-Agnostic Framework for Features at Any Resolution

ICLR 2024poster

Deep features are a cornerstone of computer vision research, capturing image semantics and enabling the community to solve downstream tasks even in the zero- or few-shot regime. However, these features often lack the spatial resolution to directly perform dense prediction tasks like segmentation and…

2022

Structure and Motion from Casual Videos

ECCV 2022poster

"Casual videos, such as those captured in daily life using a hand-held cell phone, pose problems for conventional structure-from-motion (SfM) techniques: the camera is often roughly stationary (not much parallax), and a large portion of the video may contain moving objects. Under such conditions, st…

Cited by 42SourcePDFScholar
2022

Unsupervised Semantic Segmentation by Distilling Feature Correspondences

ICLR 2022poster

Unsupervised semantic segmentation aims to discover and localize semantically meaningful categories within image corpora without any form of annotation. To solve this task, algorithms must produce features for every pixel that are both semantically meaningful and compact enough to form distinct clus…

2021

Differentiable Surface Rendering via Non-Differentiable Sampling

ICCV 2021poster

We present a method for differentiable rendering of 3D surfaces that supports both explicit and implicit representations, provides derivatives at occlusion boundaries, and is fast and simple to implement. The method first samples the surface using non-differentiable rasterization, then applies diffe…

Cited by 49PDFScholar
2021

Editing Conditional Radiance Fields

ICCV 2021poster

A neural radiance field (NeRF) is a scene model supporting high-quality view synthesis, optimized per scene. In this paper, we explore enabling user editing of a category-level NeRF trained on a shape category. Specifically, we propose a method for propagating coarse 2D user scribbles to the 3D spac…

Cited by 304PDFcodeScholar
2020

Deep Audio Priors Emerge From Harmonic Convolutional Networks

ICLR 2020poster

Convolutional neural networks (CNNs) excel in image recognition and generation. Among many efforts to explain their effectiveness, experiments show that CNNs carry strong inductive biases that capture natural image priors. Do deep networks also have inductive biases for audio signals? In this paper,…

Cited by 40SourceScholar
2018

Learning Shape Priors for Single-View 3D Completion and Reconstruction

ECCV 2018poster

The problem of single-view 3D shape completion or reconstruction is challenging, because among the many possible shapes that explain an observation, most are implausible and do not correspond to natural objects. Recent research in the field has tackled this problem by exploiting the expressiveness o…

Cited by 232SourcePDFScholar
2018

Learning to Reconstruct Shapes from Unseen Classes

NeurIPS 2018oral

From a single image, humans are able to perceive the full 3D shape of an object by exploiting learned shape priors from everyday life. Contemporary single-image 3D reconstruction algorithms aim to solve this task in a similar fashion, but often end up with priors that are highly biased by training c…

Cited by 184SourcePDFScholar
2018

Pix3D: Dataset and Methods for Single-Image 3D Shape Modeling

CVPR 2018poster

We study 3D shape modeling from a single image and make contributions to it in three aspects. First, we present Pix3D, a large-scale benchmark of diverse image-shape pairs with pixel-level 2D-3D alignment. Pix3D has wide applications in shape-related tasks including reconstruction, retrieval, viewpo…

Cited by 590SourcePDFScholar
2018

Seeing Tree Structure from Vibration

ECCV 2018poster

Humans recognize object structure from both their appearance and motion; often, motion helps to resolve ambiguities in object structure that arise when we observe object appearance only. There are particular scenarios, however, where neither appearance nor spatial-temporal motion signals are informa…

Cited by 14SourcePDFScholar
2018

Visual Object Networks: Image Generation with Disentangled 3D Representations

NeurIPS 2018poster

Recent progress in deep generative models has led to tremendous breakthroughs in image generation. While being able to synthesize photorealistic images, existing models lack an understanding of our underlying 3D world. Different from previous works built on 2D datasets and models, we present a new g…

2017

Generative Modeling of Audible Shapes for Object Perception

ICCV 2017poster

Humans infer rich knowledge of objects from both auditory and visual cues. Building a machine of such competency, however, is very challenging, due to the great difficulty in capturing large-scale, clean data of objects with both their appearance and the sound they make. In this paper, we present a…

Cited by 44PDFScholar
2017

Shape and Material from Sound

NeurIPS 2017spotlight

Hearing an object falling onto the ground, humans can recover rich information including its rough shape, material, and falling height. In this paper, we build machines to approximate such competency. We first mimic human knowledge of the physical world by building an efficient, physics-based simula…

Cited by 36SourcePDFScholar