← Search

David F. Fouhey

26 accepted papers

2025

3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination

CVPR 2025poster

The integration of language and 3D perception is crucial for embodied agents and robots that comprehend and interact with the physical world. While large language models (LLMs) have demonstrated impressive language understanding and generation capabilities, their adaptation to 3D environments (3D-LL…

2025

3D-MVP: 3D Multiview Pretraining for Manipulation

CVPR 2025poster

Recent works have shown that visual pretraining on egocentric datasets using masked autoencoders (MAE) can improve generalization for downstream robotics tasks. However, these approaches pretrain only on 2D images, while many robotics applications require 3D scene understanding. In this work, we pro…

Cited by 0SourcePDFScholar
2025

Dynamic Camera Poses and Where to Find Them

CVPR 2025poster

Annotating camera poses on dynamic Internet videos at scale is critical for advancing fields like realistic video generation and simulation. However, collecting such a dataset is difficult, as most Internet videos are unsuitable for pose estimation. Furthermore, annotating dynamic Internet videos pr…

Cited by 0SourcePDFScholar
2024

FAR: Flexible Accurate and Robust 6DoF Relative Camera Pose Estimation

CVPR 2024highlight

Estimating relative camera poses between images has been a central problem in computer vision. Methods that find correspondences and solve for the fundamental matrix offer high precision in most cases. Conversely methods predicting pose directly using neural networks are more robust to limited overl…

Cited by 4SourcePDFScholar
2024

LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent

ICRA 2024poster

3D visual grounding is a critical skill for household robots, enabling them to navigate, manipulate objects, and answer questions based on their environment. While existing approaches often rely on extensive labeled data or exhibit limitations in handling complex language queries, we propose LLM-Gro…

Cited by 100SourcecodeScholar
2023

Learning To Predict Scene-Level Implicit 3D From Posed RGBD Data

CVPR 2023poster

We introduce a method that can learn to predict scene-level implicit functions for 3D reconstruction from posed RGBD data. At test time, our system maps a previously unseen RGB image to a 3D reconstruction of a scene via implicit functions. While implicit functions for 3D reconstruction have often b…

Cited by 2SourcePDFScholar
2023

Perspective Fields for Single Image Camera Calibration

CVPR 2023highlight

Geometric camera calibration is often required for applications that understand the perspective of the image. We propose perspective fields as a representation that models the local perspective properties of an image. Perspective Fields contain per-pixel information about the camera view, parameteri…

2022

PlaneFormers: From Sparse View Planes to 3D Reconstruction

ECCV 2022poster

"We present an approach for the planar surface reconstruction of a scene from images with limited overlap. This reconstruction task is challenging since it requires jointly reasoning about single image 3D reconstruction, correspondence between images, and the relative camera pose between images. Pas…

2022

Sound Localization by Self-Supervised Time Delay Estimation

ECCV 2022poster

"Sounds reach one microphone in a stereo pair sooner than the other, resulting in an interaural time delay that conveys their directions. Estimating a sound’s time delay requires finding correspondences between the signals recorded by each microphone. We propose to learn these correspondences throug…

2022

Understanding 3D Object Articulation in Internet Videos

CVPR 2022poster

We propose to investigate detecting and characterizing the 3D planar articulation of objects from ordinary RGB videos. While seemingly easy for humans, this problem poses many challenges for computers. Our approach is based on a top-down detection system that finds planes that can be articulated. Th…

Cited by 24PDFcodeScholar
2021

PixelSynth: Generating a 3D-Consistent Experience From a Single Image

ICCV 2021poster

Recent advancements in differentiable rendering and 3D reasoning have driven exciting results in novel view synthesis from a single image. Despite realistic results, methods are limited to relatively small view change. In order to synthesize immersive scenes, models must also be able to extrapolate.…

Cited by 88PDFcodeScholar
2020

Image Transformation and CNNs: A Strategy for Encoding Human Locomotor Intent for Autonomous Wearable Robots

RA-L 2020

Wearable robots have the potential to improve the lives of countless individuals; however, challenges associated with controlling these systems must be addressed before they can reach their full potential. Modern control strategies for wearable robots are predicated on activity-specific implementati

Cited by 25SourceScholar
2020

Novel Object Viewpoint Estimation Through Reconstruction Alignment

CVPR 2020poster

The goal of this paper is to estimate the viewpoint for a novel object. Standard viewpoint estimation approaches generally fail on this task due to their reliance on a 3D model for alignment or large amounts of class-specific training data and their corresponding canonical pose. We overcome those li…

Cited by 18PDFcodeScholar
2018

Factoring Shape, Pose, and Layout From the 2D Image of a 3D Scene

CVPR 2018poster

The goal of this paper is to take a single 2D image of a scene and recover the 3D structure in terms of a small set of factors: a layout representing the enclosing surfaces as well as a set of objects represented in terms of shape and pose. We propose a convolutional neural network-based approach to…

Cited by 157SourcePDFScholar