← Search

James M. Rehg

66 accepted papers

2026

Forecasting 3D Scanpaths in Egocentric Video

CVPR 2026

Forecasting gaze behavior is an important task for understanding user intent and creating AR/VR systems that can anticipate where users will look and interact next. While prior works have addressed predicting scanpaths in static images, forecasting gaze in egocentric videos presents new challenges d

Cited by 0SourcecodeScholar
2026

Omni-MMSI: Toward Identity-attributed Social Interaction Understanding

CVPR 2026

We introduce Omni-MMSI, a new task that requires comprehensive social interaction understanding from raw audio, vision, and speech input. The task involves perceiving identity-attributed social cues (e.g., who is speaking what) and reasoning about the social interaction (e.g., whom the speaker refer

Cited by 0SourcecodeScholar
2026

Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective

CVPR 2026

Latent diffusion has become the default paradigm for visual generation, yet we observe a persistent reconstruction-generation trade-off as latent dimensionality increases: higher-capacity autoencoders improve reconstruction fidelity but generation quality eventually declines. We trace this gap to th

Cited by 0SourceScholar
2025

Gaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoders

CVPR 2025highlight

We address the problem of gaze target estimation, which aims to predict where a person is looking in a scene. Predicting a person's gaze target requires reasoning both about the person's appearance and the contents of the scene. Prior works have developed increasingly complex, hand-crafted pipelines…

2025

Improving Personalized Search with Regularized Low-Rank Parameter Updates

CVPR 2025highlight

Personalized vision-language retrieval seeks to recognize new concepts (e.g. "my dog Fido") from only a few examples. This task is challenging because it requires not only learning a new concept from a few images, but also integrating the personal and general knowledge together to recognize the conc…

2025

SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images

CVPR 2025poster

We study the problem of single-image 3D object reconstruction. Recent works have diverged into two directions: regression-based modeling and generative modeling. Regression methods efficiently infer visible surfaces, but struggle with occluded regions. Generative methods handle uncertain regions bet…

2025

ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models

CVPR 2025poster

Current diffusion-based text-to-video methods are limited to producing short video clips of a single shot and lack the capability to generate multi-shot videos with discrete transitions where the same character performs distinct activities across the same or different backgrounds. To address this li…

2025

SocialGesture: Delving into Multi-person Gesture Understanding

CVPR 2025poster

Previous research in human gesture recognition has largely overlooked multi-person interactions, which are crucial for understanding the social context of naturally occurring gestures. This limitation in existing datasets presents a significant challenge in aligning human gestures with other modalit…

Cited by 0SourcePDFScholar
2025

Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generation

CVPR 2025highlight

Symmetry is a ubiquitous and fundamental property in the visual world, serving as a critical cue for perception and structure interpretation. This paper investigates the detection of 3D reflection symmetry from a single RGB image, and reveals its significant benefit on single-image 3D generation. We…

Cited by 0SourcePDFScholar
2025

Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation

CVPR 2025highlight

Text-guided image manipulation has experienced notable advancement in recent years. In order to mitigate linguistic ambiguity, few-shot learning with visual examples has been applied for instructions that are underrepresented in the training set, or difficult to describe purely in language. However,…

Cited by 3SourcePDFScholar
2024

3x2: 3D Object Part Segmentation by 2D Semantic Correspondences

ECCV 2024poster

"3D object part segmentation is essential in computer vision applications. While substantial progress has been made in 2D object part segmentation, the 3D counterpart has received less attention, in part due to the scarcity of annotated 3D datasets, which are expensive to collect. In this work, we p…

2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning

ECCV 2024oral

"Generating instructional images of human daily actions from an egocentric viewpoint serves as a key step towards efficient skill transfer. In this paper, we introduce a novel problem – egocentric action frame generation. The goal is to synthesize an image depicting an action in the user’s context (…

2024

LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs

CVPR 2024poster

Autonomous driving (AD) has made significant strides in recent years. However existing frameworks struggle to interpret and execute spontaneous user instructions such as "overtake the car ahead." Large Language Models (LLMs) have demonstrated impressive reasoning capabilities showing potential to br…

2024

Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation

ECCV 2024poster

"Egocentric gaze anticipation serves as a key building block for the emerging capability of Augmented Reality. Notably, gaze behavior is driven by both visual cues and audio signals during daily activities. Motivated by this observation, we introduce the first model that leverages both the video and…

2024

MAPLM: A Real-World Large-Scale Vision-Language Benchmark for Map and Traffic Scene Understanding

CVPR 2024poster

Vision-language generative AI has demonstrated remarkable promise for empowering cross-modal scene understanding of autonomous driving and high-definition (HD) map systems. However current benchmark datasets lack multi-modal point cloud image and language data pairs. Recent approaches utilize visual…

2024

Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations

CVPR 2024poster

Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However most prior works on multimodal social cues focus predominantly on single-person behaviors or rely on holistic visual representations that are not aligned t…

2024

PointInfinity: Resolution-Invariant Point Diffusion Models

CVPR 2024poster

We present PointInfinity an efficient family of point cloud diffusion models. Our core idea is to use a transformer-based architecture with a fixed-size resolution-invariant latent representation. This enables efficient training with low-resolution point clouds while allowing high-resolution point c…

Cited by 10SourcePDFScholar
2024

RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models

CVPR 2024highlight

Recent advancements in diffusion-based models have demonstrated significant success in generating images from text. However video editing models have not yet reached the same level of visual quality and user control. To address this we introduce RAVE a zero-shot video editing method that leverages p…

2024

The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective

CVPR 2024poster

In recent years the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions where both visual and audio signals play a crucial role. While most prior work focus on learning about behaviors that directly involve the…

2024

ZeroShape: Regression-based Zero-shot Shape Reconstruction

CVPR 2024poster

We study the problem of single-image zero-shot 3D shape reconstruction. Recent works learn zero-shot shape reconstruction through generative modeling of 3D assets but these models are computationally expensive at train and inference time. In contrast the traditional approach to this problem is regre…

2023

Egocentric Auditory Attention Localization in Conversations

CVPR 2023poster

In a noisy conversation environment such as a dinner party, people often exhibit selective auditory attention, or the ability to focus on a particular speaker while tuning out others. Recognizing who somebody is listening to in a conversation is essential for developing technologies that can underst…

2023

ShapeClipper: Scalable 3D Shape Learning From Single-View Images via Geometric and CLIP-Based Consistency

CVPR 2023poster

We present ShapeClipper, a novel method that reconstructs 3D object shapes from real-world single-view RGB images. Instead of relying on laborious 3D, multi-view or camera pose annotation, ShapeClipper learns shape reconstruction from a set of single-view segmented images. The key idea is to facilit…

Cited by 22SourcePDFScholar
2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Egocentric Activity Recognition and Localization on a 3D Map

ECCV 2022poster

"Given a video captured from a first person perspective and the environment context of where the video is recorded, can we recognize what the person is doing and identify where the action occurs in the 3D space? We address this challenging problem of jointly recognizing and localizing actions of a m…

Cited by 27SourcePDFScholar
2022

Generative Adversarial Network for Future Hand Segmentation from Egocentric Video

ECCV 2022poster

"We introduce the novel problem of anticipating a time series of future hand masks from egocentric video. A key challenge is to model the stochasticity of future head motions, which globally impact the head-worn camera video analysis. To this end, we propose a novel deep generative model -- EgoGAN,…

2022

Planes vs. Chairs: Category-Guided 3D Shape Learning without Any 3D Cues

ECCV 2022poster

"We present a novel 3D shape reconstruction method which learns to predict an implicit 3D shape representation from a single RGB image. Our approach uses a set of single-view images of multiple object categories without viewpoint annotation, forcing the model to learn across multiple object categori…

Cited by 16SourcePDFScholar
2021

Approximate Inverse Reinforcement Learning from Vision-based Imitation Learning

ICRA 2021poster

In this work, we present a method for obtaining an implicit objective function for vision-based navigation. The proposed methodology relies on Imitation Learning, Model Predictive Control (MPC), and an interpretation technique used in Deep Neural Networks. We use Imitation Learning as a means to do…

Cited by 18SourceScholar
2021

Discriminative Appearance Modeling With Multi-Track Pooling for Real-Time Multi-Object Tracking

CVPR 2021poster

In multi-object tracking, the tracker maintains in its memory the appearance and motion information for each object in the scene. This memory is utilized for finding matches between tracks and detections, and is updated based on the matching. Many approaches model each target in isolation and lack t…

Cited by 96PDFcodeScholar
2020

A Robust Functional EM Algorithm for Incomplete Panel Count Data

NeurIPS 2020poster

Panel count data describes aggregated counts of recurrent events observed at discrete time points. To understand dynamics of health behaviors and predict future negative events, the field of quantitative behavioral research has evolved to increasingly rely upon panel count data collected via multip…

Cited by 2SourcePDFScholar
2020

Forecasting Human-Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video

ECCV 2020poster

We address the challenging task of anticipating human-object interaction in first person videos. Most existing methods either ignore how the camera wearer interacts with objects, or simply considers body motion as a separate modality. In contrast, we observe that the intentional hand movement reveal…

2020

Regularizing Neural Networks via Minimizing Hyperspherical Energy

CVPR 2020poster

Inspired by the Thomson problem in physics where the distribution of multiple propelling electrons on a unit sphere can be modeled via minimizing some potential energy, hyperspherical energy minimization has demonstrated its potential in regularizing neural networks and improving their generalizatio…

Cited by 34PDFScholar
2019

Incremental Object Learning From Contiguous Views

CVPR 2019oral

In this work, we present CRIB (Continual Recognition Inspired by Babies), a synthetic incremental object learning environment that can produce data that models visual imagery produced by object exploration in early infancy. CRIB is coupled with a new 3D object dataset, Toys-200, that contains 200 un…

Cited by 53PDFScholar
2019

Learning to Generate Synthetic Data via Compositing

CVPR 2019poster

We present a task-specific approach to synthetic data generation. Our framework employs a trainable synthesizer network that is optimized to produce meaningful training samples by assessing the strengths and weaknesses of a 'target' classifier. The synthesizer and target networks are trained in an a…

Cited by 169PDFScholar
2019

Locally Weighted Regression Pseudo-Rehearsal for Adaptive Model Predictive Control

CoRL 2019

We consider the problem of online adaptation of a neural network designed to represent system dynamics. The neural network model is intended to be used by an MPC control law for autonomous control. This problem is challenging because both input and target distributions are non-stationary, and naive

Cited by 0SourcePDFScholar
2019

Taking a Deeper Look at the Inverse Compositional Algorithm

CVPR 2019oral

In this paper, we provide a modern synthesis of the classic inverse compositional algorithm for dense image alignment. We first discuss the assumptions made by this well-established technique, and subsequently propose to relax these assumptions by incorporating data-driven priors into this model. Mo…

Cited by 63PDFcodeScholar
2019

Unsupervised 3D Pose Estimation With Geometric Self-Supervision

CVPR 2019poster

We present an unsupervised learning approach to re- cover 3D human pose from 2D skeletal joints extracted from a single image. Our method does not require any multi- view image data, 3D skeletons, correspondences between 2D-3D points, or use previously learned 3D priors during training. A lifting ne…

Cited by 249PDFScholar
2019

Vision-Based High-Speed Driving With a Deep Dynamic Observer

RA-L 2019

In this letter, we present a framework for combining deep learning-based road detection, particle filters, and model predictive control (MPC) to drive aggressively using only a monocular camera, IMU, and wheel speed sensors. This framework uses deep convolutional neural networks combined with LSTMs

Cited by 52SourceScholar
2018

Best Response Model Predictive Control for Agile Interactions Between Autonomous Ground Vehicles

ICRA 2018poster

We introduce an algorithm for autonomous control of multiple fast ground vehicles operating in close proximity to each other. The algorithm is based on a combination of the game theoretic notion of iterated best response, and an information theoretic model predictive control algorithm designed for n…

Cited by 65SourceScholar
2018

Connecting Gaze, Scene, and Attention: Generalized Attention Estimation via Joint Modeling of Gaze and Scene Saliency

ECCV 2018poster

This paper addresses the challenging problem of estimating the general visual attention of people in images. Our proposed method is designed to work across multiple naturalistic social scenarios and provides a full picture of the subject’s attention and gaze. In contrast, earlier works on gaze and a…

2018

In the Eye of Beholder: Joint Learning of Gaze and Actions in First Person Video

ECCV 2018poster

We address the task of jointly determining what a person is doing and where they are looking based on the analysis of video captured by a headworn camera. We propose a novel deep model for joint gaze estimation and action recognition in First Person Vision. Our method describes the participant's gaz…

Cited by 411SourcePDFScholar
2018

Learning Rigidity in Dynamic Scenes with a Moving Camera for 3D Motion Field Estimation

ECCV 2018poster

Estimation of 3D motion in a dynamic scene from a temporal pair of images is a core task in many scene understanding problems. In real world applications, a dynamic scene is commonly captured by a moving camera (i.e., panning, tilting or hand-held), increasing the task complexity because the scene i…

2017

Aggressive Deep Driving: Combining Convolutional Neural Networks and Model Predictive Control

CoRL 2017

We present a framework for vision-based model predictive control (MPC) for the task of aggressive, high-speed autonomous driving. Our approach uses deep convolutional neural networks to predict cost functions from input video which are directly suitable for online trajectory optimization with MPC. W

Cited by 0SourcePDFScholar
2017

Information theoretic MPC for model-based reinforcement learning

ICRA 2017poster

We introduce an information theoretic model predictive control (MPC) algorithm capable of handling complex cost criteria and general nonlinear dynamics. The generality of the approach makes it possible to use multi-layer neural networks as dynamics models, which we incorporate into our MPC algorithm…

Cited by 711SourceScholar
2017

iSurvive: An Interpretable, Event-time Prediction Model for mHealth

ICML 2017poster

An important mobile health (mHealth) task is the use of multimodal data, such as sensor streams and self-report, to construct interpretable time-to-event predictions of, for example, lapse to alcohol or illicit drug use. Interpretability of the prediction model is important for acceptance and adopti…

Cited by 30SourcePDFScholar
2016

Aggressive driving with model predictive path integral control

ICRA 2016

In this paper we present a model predictive control algorithm designed for optimizing non-linear systems subject to complex cost criteria. The algorithm is based on a stochastic optimal control framework using a fundamental relationship between the information theoretic notions of free energy and re

Cited by 578SourceScholar
2015

Combining tactile sensing and vision for rapid haptic mapping

IROS 2015poster

We consider the problem of enabling a robot to efficiently obtain a dense haptic map of its visible surroundings using the complementary properties of vision and tactile sensing. Our approach assumes that visible surfaces that look similar to one another are likely to have similar haptic properties.…

Cited by 27SourceScholar
2015

Efficient Learning of Continuous-Time Hidden Markov Models for Disease Progression

NeurIPS 2015poster

The Continuous-Time Hidden Markov Model (CT-HMM) is an attractive approach to modeling disease progression due to its ability to describe noisy observations arriving irregularly in time. However, the lack of an efficient parameter learning algorithm for CT-HMM restricts its use to very small models…

Cited by 146SourcePDFScholar
2015

Gaze-Enabled Egocentric Video Summarization via Constrained Submodular Maximization

CVPR 2015poster

With the proliferation of wearable cameras, the number of videos of users documenting their personal lives using such devices is rapidly increasing. Since such videos may span hours, there is an important need for mechanisms that represent the information content in a compact form (i.e., shorter…

Cited by 207SourcePDFScholar
2015

Minimizing Human Effort in Interactive Tracking by Incremental Learning of Model Parameters

ICCV 2015poster

We address the problem of minimizing human effort in interactive tracking by learning sequence-specific model parameters. Determining the optimal model parameters for each sequence is a critical problem in tracking. We demonstrate that by using the optimal model parameters for each sequence we can a…

Cited by 6PDFScholar
2015

Multi-scale perception and path planning on probabilistic obstacle maps

ICRA 2015poster

We present a path-planning algorithm that leverages a multi-scale representation of the environment. The algorithm works in n dimensions. The information of the environment is stored in a tree representing a recursive dyadic partitioning of the search space. The information used by the algorithm is…

Cited by 22SourceScholar
2015

Robust Video Segment Proposals With Painless Occlusion Handling

CVPR 2015poster

We propose a robust algorithm to generate video segment proposals. The proposals generated by our method can start from any frame in the video and are robust to complete occlusions. Our method does not assume specific motion models and even has a limited capability to generalize across videos. We bu…

Cited by 35SourcePDFScholar
2015

The Middle Child Problem: Revisiting Parametric Min-Cut and Seeds for Object Proposals

ICCV 2015poster

Object proposals have recently fueled the progress in detection performance. These proposals aim to provide category-agnostic localizations for all objects in an image. One way to generate proposals is to perform parametric min-cuts over seed locations. This paper demonstrates that standard parametr…

Cited by 31PDFScholar