← Search

Bryan Russell

27 accepted papers

2025

Discovering Divergent Representations between Text-to-Image Models

ICCV 2025poster

In this paper, we investigate when and how visual representations learned by two different generative models diverge from each other. Specifically, given two text-to-image models, our goal is to discover visual attributes that appear in images generated by one model but not the other, along with the…

Cited by 0SourcePDFScholar
2025

Improving Personalized Search with Regularized Low-Rank Parameter Updates

CVPR 2025highlight

Personalized vision-language retrieval seeks to recognize new concepts (e.g. "my dog Fido") from only a few examples. This task is challenging because it requires not only learning a new concept from a few images, but also integrating the personal and general knowledge together to recognize the conc…

2025

ResidualViT for Efficient Temporally Dense Video Encoding

ICCV 2025poster

Several video understanding tasks, such as natural language temporal video grounding, temporal activity localization, and audio description generation, require "temporally dense" reasoning over frames sampled at high temporal resolution. However, computing frame-level features for these tasks is com…

Cited by 0SourcePDFScholar
2025

Video-Guided Foley Sound Generation with Multimodal Controls

CVPR 2025poster

Generating sound effects for videos often requires creating artistic sound effects that diverge significantly from real-life sources and flexible control in the sound design. To address this problem, we introduce *MultiFoley*, a model designed for video-guided sound generation that supports multimod…

Cited by 10SourcePDFScholar
2024

Koala: Key Frame-Conditioned Long Video-LLM

CVPR 2024highlight

Long video question answering is a challenging task that involves recognizing short-term activities and reasoning about their fine-grained relationships. State-of-the-art video Large Language Models (vLLMs) hold promise as a viable solution due to their demonstrated emergent capabilities on new task…

Cited by 33SourcePDFScholar
2023

Conditional Generation of Audio From Video via Foley Analogies

CVPR 2023poster

The sound effects that designers add to videos are designed to convey a particular artistic effect and, thus, may be quite different from a scene's true sound. Inspired by the challenges of creating a soundtrack for a video that differs from its true sound, but that nonetheless matches the actions o…

2023

Language-Guided Audio-Visual Source Separation via Trimodal Consistency

CVPR 2023poster

We propose a self-supervised approach for learning to perform audio source separation in videos based on natural language queries, using only unlabeled video and audio pairs as training data. A key challenge in this task is learning to associate the linguistic description of a sound-emitting object…

Cited by 19SourcePDFScholar
2023

Language-Guided Music Recommendation for Video via Prompt Analogies

CVPR 2023highlight

We propose a method to recommend music for an input video while allowing a user to guide music selection with free-form natural language. A key challenge of this problem setting is that existing music video datasets provide the needed (video, music) training pairs, but lack text descriptions of the…

Cited by 30SourcePDFScholar
2023

Meta-Personalizing Vision-Language Models To Find Named Instances in Video

CVPR 2023poster

Large-scale vision-language models (VLM) have shown impressive results for language-guided search applications. While these models allow category-level queries, they currently struggle with personalized searches for moments in a video where a specific object instance such as "My dog Biscuit" appears…

2022

Focal Length and Object Pose Estimation via Render and Compare

CVPR 2022poster

We introduce FocalPose, a neural render-and-compare method for jointly estimating the camera-object 6D pose and camera focal length given a single RGB input image depicting a known object. The contributions of this work are twofold. First, we derive a focal length update rule that extends an existin…

Cited by 25PDFcodeScholar
2022

Monocular Dynamic View Synthesis: A Reality Check

NeurIPS 2022accept

We study the recent progress on dynamic view synthesis (DVS) from monocular video. Though existing approaches have demonstrated impressive results, we show a discrepancy between the practical capture process and the existing experimental protocols, which effectively leaks in multi-view signals durin…

2021

Editing Conditional Radiance Fields

ICCV 2021poster

A neural radiance field (NeRF) is a scene model supporting high-quality view synthesis, optimized per scene. In this paper, we explore enabling user editing of a category-level NeRF trained on a shape category. Specifically, we propose a method for propagating coarse 2D user scribbles to the 3D spac…

Cited by 300PDFcodeScholar
2021

Look at What I’m Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

NeurIPS 2021spotlight

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision on a large corpus of videos with accompanying transcribed narrations. To achieve this goal, we propose a multilayer cro…

Cited by 26SourcePDFScholar
2021

Weakly Supervised Human-Object Interaction Detection in Video via Contrastive Spatiotemporal Regions

ICCV 2021poster

We introduce the task of weakly supervised learning for detecting human and object interactions in videos. Our task poses unique challenges as a system does not know what types of human-object interactions are present in a video or the actual spatiotemporal location of the human and object. To addre…

Cited by 13PDFcodeScholar
2020

Contact and Human Dynamics from Monocular Video

ECCV 2020poster

Existing deep models predict 2D and 3D kinematic poses from video that are approximately accurate, but contain visible errors that violate physical constraints, such as feet penetrating the ground and bodies leaning at extreme angles. In this paper, we present a physics-based method for inferring 3D…

2019

Bounce and Learn: Modeling Scene Dynamics with Real-World Bounces

ICLR 2019poster

We introduce an approach to model surface properties governing bounces in everyday scenes. Our model learns end-to-end, starting from sensor inputs, to predict post-bounce trajectories and infer two underlying physical properties that govern bouncing - restitution and effective collision normals. O…

Cited by 26SourcePDFScholar
2019

FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB Images

ICCV 2019poster

Estimating 3D hand pose from single RGB images is a highly ambiguous problem that relies on an unbiased training dataset. In this paper, we analyze cross-dataset generalization when training on existing datasets. We find that approaches perform well on the datasets they are trained on, but do not ge…

Cited by 544PDFScholar
2019

Learning elementary structures for 3D shape generation and matching

NeurIPS 2019poster

We propose to represent shapes as the deformation and combination of learnt elementary 3D structures. We demonstrate this decomposition in learnt elementary 3D structures is highly interpretable and leads to clear improvements in 3D shape generation and matching. More precisely, we present two comp…

2019

Neural Re-Simulation for Generating Bounces in Single Images

ICCV 2019poster

We introduce a method to generate videos of dynamic virtual objects plausibly interacting via collisions with a still image's environment. Given a starting trajectory, physically simulated with the estimated geometry of a single, static input image, we learn to 'correct' this trajectory to a visuall…

Cited by 12PDFScholar
2018

BodyNet: Volumetric Inference of 3D Human Body Shapes

ECCV 2018poster

Human shape estimation is an important task for video editing, animation and fashion industry. Predicting 3D human body shape from natural images, however, is highly challenging due to factors such as variation in human bodies, clothing and viewpoint. Prior methods addressing this problem typically…

Cited by 530SourcePDFScholar
2017

ActionVLAD: Learning Spatio-Temporal Aggregation for Action Classification

CVPR 2017poster

In this work, we introduce a new video representation for action classification that aggregates local convolutional features across the entire spatio-temporal extent of the video. We do so by integrating state-of-the-art two-stream networks with learnable spatio-temporal feature aggregation. The res…

Cited by 607PDFScholar
2017

Localizing Moments in Video With Natural Language

ICCV 2017poster

We consider retrieving a specific temporal segment, or moment, from a video given a natural language text description. Methods designed to retrieve whole video clips with natural language determine what occurs in a video but not when. To address this issue, we propose the Moment Context Network (MCN…

Cited by 1155PDFScholar
2016

SURGE: Surface Regularized Geometry Estimation from a Single Image

NeurIPS 2016poster

This paper introduces an approach to regularize 2.5D surface normal and depth predictions at each pixel given a single input image. The approach infers and reasons about the underlying 3D planar surfaces depicted in the image to snap predicted normals and depths to inferred planar surfaces, all whil…

Cited by 103SourcePDFScholar