← Search

James Matthew Rehg

13 accepted papers

2026

DiffVax: Optimization-Free Image Immunization Against Diffusion-Based Editing

ICLR 2026poster

Current image immunization defense techniques against diffusion-based editing embed imperceptible noise into target images to disrupt editing models. However, these methods face scalability challenges, as they require time-consuming optimization for each image separately, taking hours for small batc…

Cited by 0SourcecodeScholar
2025

Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation

NeurIPS 2025spotlight

Humans and traditional computer vision methods rely on a diverse set of monocular cues to infer 3D structure from a single image, such as shading, texture, silhouette, etc. While recent deep generative models have dramatically advanced single-image 3D generation, it remains unclear which image cues…

Cited by 0SourceScholar
2025

DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural Images

NeurIPS 2025poster

Numerous models have been developed for scanpath and saliency prediction, which are typically trained on scanpaths, which model eye movement as a sequence of discrete fixation points connected by saccades, while the rich information contained in the raw trajectories is often discarded. Moreover, mos…

Cited by 0SourcecodeScholar
2025

Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs

NeurIPS 2025poster

Current Vision-Language Models (VLMs) struggle with fine-grained spatial reasoning, particularly when multi-step logic and precise spatial alignment are required. In this work, we introduce SpatialReasoner-R1, a vision-language reasoning model designed to address these limitations. To construct high…

Cited by 0SourceScholar
2025

RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data

ICLR 2025poster

We present RelCon, a novel self-supervised Relative Contrastive learning approach for training a motion foundation model from wearable accelerometry sensors. First, a learnable distance measure is trained to capture motif similarity and domain-specific semantic information such as rotation invarianc…

2025

Toward Human Deictic Gesture Target Estimation

NeurIPS 2025poster

Humans have a remarkable ability to use co-speech deictic gestures, such as pointing and showing, to enrich verbal communication and support social interaction. These gestures are so fundamental that infants begin to use them even before they acquire spoken language, which highlights their central r…

Cited by 0SourcecodeScholar
2024

REBAR: Retrieval-Based Reconstruction for Time-series Contrastive Learning

ICLR 2024poster

The success of self-supervised contrastive learning hinges on identifying positive data pairs, such that when they are pushed together in embedding space, the space encodes useful information for subsequent downstream tasks. Constructing positive pairs is non-trivial as the pairing must be similar e…

2023

Low-shot Object Learning with Mutual Exclusivity Bias

NeurIPS 2023poster

This paper introduces Low-shot Object Learning with Mutual Exclusivity Bias (LSME), the first computational framing of mutual exclusivity bias, a phenomenon commonly observed in infants during word learning. We provide a novel dataset, comprehensive baselines, and a SOTA method to enable the ML comm…

2022

Kernel Multimodal Continuous Attention

NeurIPS 2022accept

Attention mechanisms take an expectation of a data representation with respect to probability weights. Recently, (Martins et al. 2020, 2021) proposed continuous attention mechanisms, focusing on unimodal attention densities from the exponential and deformed exponential families: the latter has spars…

Cited by 2SourcePDFScholar
2022

Learning Dense Object Descriptors from Multiple Views for Low-shot Category Generalization

NeurIPS 2022accept

A hallmark of the deep learning era for computer vision is the successful use of large-scale labeled datasets to train feature representations. This has been done for tasks ranging from object recognition and semantic segmentation to optical flow estimation and novel view synthesis of 3D scenes. In…

2022

PulseImpute: A Novel Benchmark Task for Pulsative Physiological Signal Imputation

NeurIPS 2022accept

The promise of Mobile Health (mHealth) is the ability to use wearable sensors to monitor participant physiology at high frequencies during daily life to enable temporally-precise health interventions. However, a major challenge is frequent missing data. Despite a rich imputation literature, existing…

2021

No RL, No Simulation: Learning to Navigate without Navigating

NeurIPS 2021poster

Most prior methods for learning navigation policies require access to simulation environments, as they need online policy interaction and rely on ground-truth maps for rewards. However, building simulators is expensive (requires manual effort for each and every scene) and creates challenges in trans…

Cited by 93SourcePDFScholar