← Search

David W. Jacobs

17 accepted papers

2024

CALVIN: Improved Contextual Video Captioning via Instruction Tuning

NeurIPS 2024poster

The recent emergence of powerful Vision-Language models (VLMs) has significantly improved image captioning. Some of these models are extended to caption videos as well. However, their capabilities to understand complex scenes are limited, and the descriptions they provide for scenes tend to be overl…

Cited by 0SourcePDFScholar
2024

Rethinking Score Distillation as a Bridge Between Image Distributions

NeurIPS 2024poster

Score distillation sampling (SDS) has proven to be an important tool, enabling the use of large-scale diffusion priors for tasks operating in data-poor domains. Unfortunately, SDS has a number of characteristic artifacts that limit its utility in general-purpose applications. In this paper, we make…

Cited by 12SourcePDFScholar
2022

Autoregressive Perturbations for Data Poisoning

NeurIPS 2022accept

The prevalence of data scraping from social media as a means to obtain datasets has led to growing concerns regarding unauthorized use of data. Data poisoning attacks have been proposed as a bulwark against scraping, as they make data ``unlearnable'' by adding small, imperceptible perturbations. Unf…

2019

Neural Inverse Rendering of an Indoor Scene From a Single Image

ICCV 2019poster

Inverse rendering aims to estimate physical attributes of a scene, e.g., reflectance, geometry, and lighting, from image(s). Inverse rendering has been studied primarily for single objects or with methods that solve for only one of the scene attributes. We propose the first learning based approach t…

Cited by 164PDFScholar
2018

End-to-End Recovery of Human Shape and Pose

CVPR 2018poster

We describe Human Mesh Recovery (HMR), an end-to-end framework for reconstructing a full 3D mesh of a human body from a single RGB image. In contrast to most current methods that compute 2D or 3D joint locations, we produce a richer and more useful mesh representation that is parameterized by shape…

2018

Label Denoising Adversarial Network (LDAN) for Inverse Lighting of Faces

CVPR 2018poster

Lighting estimation from faces is an important task and has applications in many areas such as image editing, intrinsic image decomposition, and image forgery detection. We propose to train a deep Convolutional Neural Network (CNN) to regress lighting parameters from a single face image. Lacking mas…

Cited by 23SourcePDFScholar
2018

SfSNet: Learning Shape, Reflectance and Illuminance of Faces `in the Wild'

CVPR 2018poster

We present SfSNet, an end-to-end learning framework for producing an accurate decomposition of an unconstrained human face image into shape, reflectance and illuminance. SfSNet is designed to reflect a physical lambertian rendering model. SfSNet learns from a mixture of labeled synthetic and unlabel…

Cited by 377SourcePDFScholar
2017

3D Menagerie: Modeling the 3D Shape and Pose of Animals

CVPR 2017spotlight

There has been significant work on learning realistic, articulated, 3D models of the human body. In contrast, there are few such models of animals, despite many applications. The main challenge is that animals are much less cooperative than humans. The best human body models are learned from thousan…

Cited by 491PDFScholar
2017

A New Rank Constraint on Multi-View Fundamental Matrices, and Its Application to Camera Location Recovery

CVPR 2017spotlight

Accurate estimation of camera matrices is an important step in structure from motion algorithms. In this paper we introduce a novel rank constraint on collections of fundamental matrices in multi-view settings. We show that in general, with the selection of proper scale factors, a matrix formed by s…

Cited by 31PDFScholar