← Search

Oliver Wang

34 accepted papers

2025

Motion Prompting: Controlling Video Generation with Motion Trajectories

CVPR 2025poster

Motion control is crucial for generating expressive and compelling video content; however, most existing video generation models rely mainly on text prompts for control, which struggle to capture the nuances of dynamic actions and temporal compositions. To this end, we train a video generation model…

Cited by 22SourcePDFScholar
2024

Image Neural Field Diffusion Models

CVPR 2024highlight

Diffusion models have shown an impressive ability to model complex data distributions with several key advantages over GANs such as stable training better coverage of the training distribution's modes and the ability to solve inverse problems without extra training. However most diffusion models lea…

Cited by 6SourcePDFScholar
2024

Readout Guidance: Learning Control from Diffusion Features

CVPR 2024highlight

We present Readout Guidance a method for controlling text-to-image diffusion models with learned signals. Readout Guidance uses readout heads lightweight networks trained to extract signals from the features of a pre-trained frozen diffusion model at every timestep. These readouts can encode single-…

Cited by 26SourcePDFScholar
2024

VecFusion: Vector Font Generation with Diffusion

CVPR 2024highlight

We present VecFusion a new neural architecture that can generate vector fonts with varying topological structures and precise control point positions. Our approach is a cascaded diffusion model which consists of a raster diffusion model followed by a vector diffusion model. The raster model generate…

Cited by 9SourcePDFScholar
2023

Perspective Fields for Single Image Camera Calibration

CVPR 2023highlight

Geometric camera calibration is often required for applications that understand the perspective of the image. We propose perspective fields as a representation that models the local perspective properties of an image. Perspective Fields contain per-pixel information about the camera view, parameteri…

2021

Differentiable Signal Processing With Black-Box Audio Effects

ICASSP 2021accepted

We present a data-driven approach to automate audio signal processing by incorporating stateful third-party, audio effects as layers within a deep neural network. We then train a deep encoder to analyze input audio and control effect parameters to perform the desired signal manipulation, requiring o…

Cited by 0SourceScholar
2021

Learning To Recover 3D Scene Shape From a Single Image

CVPR 2021poster

Despite significant progress in monocular depth estimation in the wild, recent state-of-the-art methods cannot be used to recover accurate 3D scene shape due to an unknown depth shift induced by shift-invariant reconstruction losses used in mixed-data depth prediction training, and possible unknown…

Cited by 284PDFcodeScholar
2021

Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes

CVPR 2021poster

We present a method to perform novel view and time synthesis of dynamic scenes, requiring only a monocular video with known camera poses as input. To do this, we introduce Neural Scene Flow Fields, a new representation that models the dynamic scene as a time-variant continuous function of appearance…

Cited by 862PDFScholar
2021

Real-Time Semantic Segmentation With Fast Attention

RA-L 2021

In deep CNN based models for semantic segmentation, high accuracy relies on rich spatial context (large receptive fields) and fine spatial details (high resolution), both of which incur high computational costs. In this letter, we propose a novel architecture that addresses both challenges and achie

Cited by 143SourcecodeScholar
2020

CNN-Generated Images Are Surprisingly Easy to Spot... for Now

CVPR 2020oral

In this work we ask whether it is possible to create a "universal" detector for telling apart real images from these generated by a CNN, regardless of architecture or dataset used. To test this, we collect a dataset consisting of fake images generated by 11 different CNN-based image generator models…

Cited by 1237PDFcodeScholar
2020

LandscapeAR: Large Scale Outdoor Augmented Reality by Matching Photographs with Terrain Models Using Learned Descriptors

ECCV 2020poster

We introduce a solution to large scale Augmented Reality for outdoor scenes by registering camera images to textured Digital Elevation Models (DEMs). To accomodate the inherent differences in appearance between real images and DEMs, we train a cross-domain feature descriptor using Structure From Mot…

Cited by 22SourcePDFScholar
2020

SDC-Depth: Semantic Divide-and-Conquer Network for Monocular Depth Estimation

CVPR 2020poster

Monocular depth estimation is an ill-posed problem, and as such critically relies on scene priors and semantics. Due to its complexity, we propose a deep neural network model based on a semantic divide-and-conquer approach. Our model decomposes a scene into semantic segments, such as object instance…

Cited by 160PDFScholar
2020

Structure-Guided Ranking Loss for Single Image Depth Prediction

CVPR 2020poster

Single image depth prediction is a challenging task due to its ill-posed nature and challenges with capturing ground truth for supervision. Large-scale disparity data generated from stereo photos and 3D videos is a promising source of supervision, however, such disparity data can only approximate th…

Cited by 214PDFcodeScholar
2020

Swapping Autoencoder for Deep Image Manipulation

NeurIPS 2020poster

Deep generative models have become increasingly effective at producing realistic images from randomly sampled seeds, but using such models for controllable manipulation of existing images remains challenging. We propose the Swapping Autoencoder, a deep model designed specifically for image manipulat…

Cited by 403SourcePDFScholar
2020

Temporally Distributed Networks for Fast Video Semantic Segmentation

CVPR 2020poster

We present TDNet, a temporally distributed network designed for fast and accurate video semantic segmentation. We observe that features extracted from a certain high-level layer of a deep CNN can be approximated by composing features extracted from several shallower sub-networks. Leveraging the inhe…

Cited by 250PDFScholar
2019

Detecting Photoshopped Faces by Scripting Photoshop

ICCV 2019poster

Most malicious photo manipulations are created using standard image editing tools, such as Adobe Photoshop. We present a method for detecting one very popular Photoshop manipulation -- image warping applied to human faces -- using a model trained entirely using fake images that were automatically ge…

Cited by 180PDFScholar
2019

Interactive Sketch & Fill: Multiclass Sketch-to-Image Translation

ICCV 2019poster

We propose an interactive GAN-based sketch-to-image translation method that helps novice users easily create images of simple objects. The user starts with a sparse sketch and a desired object category, and the network then recommends its plausible completion(s) and shows a corresponding synthesized…

Cited by 166PDFcodeScholar
2019

Photometric Mesh Optimization for Video-Aligned 3D Object Reconstruction

CVPR 2019poster

In this paper, we address the problem of 3D object mesh reconstruction from RGB videos. Our approach combines the best of multi-view geometric and data-driven methods for 3D reconstruction by optimizing object meshes for multi-view photometric consistency while constraining mesh deformations with a…

Cited by 91PDFcodeScholar
2018

Learning Blind Video Temporal Consistency

ECCV 2018poster

Applying image processing algorithms independently to each frame of a video often leads to undesired inconsistent results over time. Developing temporally consistent video-based extensions, however, requires domain knowledge for individual tasks and is unable to generalize to other applications. In…

2018

ST-GAN: Spatial Transformer Generative Adversarial Networks for Image Compositing

CVPR 2018poster

We address the problem of finding realistic geometric corrections to a foreground object such that it appears natural when composited into a background image. To achieve this, we propose a novel Generative Adversarial Network (GAN) architecture that utilizes Spatial Transformer Networks (STNs) as th…

2018

Self-Supervised Generation of Spatial Audio for 360° Video

NeurIPS 2018poster

We introduce an approach to convert mono audio recorded by a 360° video camera into spatial audio, a representation of the distribution of sound over the full viewing sphere. Spatial audio is an important component of immersive 360° video viewing, but spatial audio microphones are still rare in curr…

Cited by 201SourcePDFScholar
2018

The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

CVPR 2018poster

While it is nearly effortless for humans to quickly assess the perceptual similarity between two images, the underlying processes are thought to be quite complex. Despite this, the most widely used perceptual metrics today, such as PSNR and SSIM, are simple, shallow functions, and fail to account f…

2017

Deep Video Deblurring for Hand-Held Cameras

CVPR 2017spotlight

Motion blur from camera shake is a major problem in videos captured by hand-held devices. Unlike single-image deblurring, video-based approaches can take advantage of the abundant information that exists across neighboring frames. As a result the best performing methods rely on the alignment of near…

Cited by 711PDFScholar
2017

High-Resolution Image Inpainting Using Multi-Scale Neural Patch Synthesis

CVPR 2017poster

Recent advances in deep learning have shown exciting promise in filling large holes in natural images with semantically plausible and context aware details, impacting fundamental image manipulation tasks such as object removal. While these learning-based methods are significantly more effective in c…

Cited by 1114PDFScholar
2017

Localizing Moments in Video With Natural Language

ICCV 2017poster

We consider retrieving a specific temporal segment, or moment, from a video given a natural language text description. Methods designed to retrieve whole video clips with natural language determine what occurs in a video but not when. To address this issue, we propose the Moment Context Network (MCN…

Cited by 1155PDFScholar
2017

Toward Multimodal Image-to-Image Translation

NeurIPS 2017poster

Many image-to-image translation problems are ambiguous, as a single input image may correspond to multiple possible outputs. In this work, we aim to model a distribution of possible outputs in a conditional generative modeling setting. The ambiguity of the mapping is distilled in a low-dimensional l…

2016

Efficient Large-Scale Approximate Nearest Neighbor Search on the GPU

CVPR 2016poster

We present a new approach for efficient approximate nearest neighbor (ANN) search in high dimensional spaces, extending the idea of Product Quantization. We propose a two level product and vector quantization tree that reduces the number of vector comparisons required during tree traversal. Our appr…

Cited by 89PDFScholar
2015

FaceDirector: Continuous Control of Facial Performance in Video

ICCV 2015poster

We present a method to continuously blend between multiple facial performances of an actor, which can contain different facial expressions or emotional states. As an example, given sad and angry video takes of a scene, our method empowers the movie director to specify arbitrary weighted combinations…

Cited by 19PDFScholar
2015

Fully Connected Object Proposals for Video Segmentation

ICCV 2015poster

We present a novel approach to video segmentation using multiple object proposals. The problem is formulated as a minimization of a novel energy function defined over a fully connected graph of object proposals. Our model combines appearance with long-range point tracks, which is key to ensure robus…

Cited by 210PDFScholar
2015

Phase-Based Frame Interpolation for Video

CVPR 2015poster

Standard approaches to computing interpolated (in-between) frames in a video sequence require accurate pixel correspondences between images e.g. using optical flow. We present an efficient alternative by leveraging recent developments in phase-based methods that represent motion in the phase shift o…

Cited by 362SourcePDFScholar
2015

Scalable Structure From Motion for Densely Sampled Videos

CVPR 2015poster

Videos consisting of thousands of high resolution frames are challenging for existing structure from motion (SfM) and simultaneous-localization and mapping (SLAM) techniques. We present a new approach for simultaneously computing extrinsic camera poses and 3D scene structure that is capable of handl…