← Search

David Fouhey

14 accepted papers

2025

Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme

NeurIPS 2025poster

Conditional diffusion models provide a natural framework for probabilistic prediction of dynamical systems and have been successfully applied to fluid dynamics and weather prediction. However, in many settings, the available information at a given time represents only a small fraction of what is nee…

Cited by 0SourceScholar
2025

Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos

CVPR 2025poster

Learning to understand dynamic 3D scenes from imagery is crucial for applications ranging from robotics to scene reconstruction. Yet, unlike other problems where large-scale supervised training has enabled rapid progress, directly supervising methods for recovering 3D motion remains challenging due…

2024

DEPICT: Diffusion-Enabled Permutation Importance for Image Classification Tasks

ECCV 2024poster

"We propose a permutation-based explanation method for image classifiers. Current image-model explanations like activation maps are limited to instance-based explanations in the pixel space, making it difficult to understand global model behavior. In contrast, permutation based explanations for tabu…

Cited by 1SourcePDFScholar
2024

Multi-Object Hallucination in Vision Language Models

NeurIPS 2024poster

Large vision language models (LVLMs) often suffer from object hallucination, producing objects not present in the given images. While current benchmarks for object hallucination primarily concentrate on the presence of a single object class rather than individual entities, this work systematically…

2024

NIFTY: Neural Object Interaction Fields for Guided Human Motion Synthesis

CVPR 2024poster

We address the problem of generating realistic 3D motions of humans interacting with objects in a scene. Our key idea is to create a neural interaction field attached to a specific object which outputs the distance to the valid interaction manifold given a human pose as input. This interaction field…

Cited by 44SourcePDFScholar
2024

Reconstructing Hands in 3D with Transformers

CVPR 2024poster

We present an approach that can reconstruct hands in 3D from monocular input. Our approach for Hand Mesh Recovery HaMeR follows a fully transformer-based architecture and can analyze hands with significantly increased accuracy and robustness compared to previous work. The key to HaMeR's success lies…

2023

EPIC Fields: Marrying 3D Geometry and Video Understanding

NeurIPS 2023poster

Neural rendering is fuelling a unification of learning, 3D geometry and video understanding that has been waiting for more than two decades. Progress, however, is still hampered by a lack of suitable datasets and benchmarks. To address this gap, we introduce EPIC Fields, an augmentation of EPIC-KITC…

2023

Towards A Richer 2D Understanding of Hands at Scale

NeurIPS 2023poster

As humans, we learn a lot about how to interact with the world by observing others interacting with their hands. To help AI systems obtain a better understanding of hand interactions, we introduce a new model that produces a rich understanding of hand interaction. Our system produces a richer output…

Cited by 17SourcePDFScholar
2022

EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations

NeurIPS 2022accept

We introduce VISOR, a new dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video. VISOR annotates videos from EPIC-KITCHENS, which comes with a new set of challenges not encountered in current video segmentation datasets. Specifically, we need…

2021

COHESIV: Contrastive Object and Hand Embedding Segmentation In Video

NeurIPS 2021poster

In this paper we learn to segment hands and hand-held objects from motion. Our system takes a single RGB image and hand location as input to segment the hand and hand-held object. For learning, we generate responsibility maps that show how well a hand's motion explains other pixels' motion in video.…

Cited by 19SourcePDFScholar
2019

Cross-Task Weakly Supervised Learning From Instructional Videos

CVPR 2019poster

In this paper we investigate learning visual models for the steps of ordinary tasks using weak supervision via instructional narrations and an ordered list of steps instead of strong supervision via temporal annotations. At the heart of our approach is the observation that weakly supervised learning…

Cited by 311PDFcodeScholar
2018

Visual Memory for Robust Path Following

NeurIPS 2018oral

Humans routinely retrace a path in a novel environment both forwards and backwards despite uncertainty in their motion. In this paper, we present an approach for doing so. Given a demonstration of a path, a first network generates an abstraction of the path. Equipped with this abstraction, a second…

Cited by 64SourcePDFScholar