← Search

Changan Chen

30 accepted papers

2026

A Pragmatist Robot: Learning to Plan Tasks by Experiencing the Real World

RA-L 2026

Large language models (LLMs) have emerged as the dominant paradigm for robotic task planning using natural language instructions. However, trained on general internet data, LLMs are not inherently aligned with the embodiment, skill sets, and limitations of real-world robotic systems. Inspired by the

Cited by 2SourcecodeScholar
2026

BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models

CVPR 2026

Autoregressive video models are promising for world modeling via next-frame prediction, but they suffer from exposure bias: a mismatch between training on clean contexts and inference on self-generated frames, causing errors to compound and quality to drift over time. We introduce Backwards Aggregat

Cited by 0SourceScholar
2026

Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation

ICML 2026poster

Latent diffusion models excel at generating high-quality images but lose the benefits of end-to-end modeling. They discard information during image encoding, require a separately trained decoder, and model an auxiliary distribution to the raw data. In this paper, we propose Latent Forcing, a simple …

Cited by 0SourceScholar
2026

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

CVPR 2026

Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human behavior as a translation task--co-speech gesture or text-to-motion that maps a fixed utterance to motion clips--without requiring agentic decision-maki

Cited by 0SourceScholar
2025

The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion

CVPR 2025poster

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for creating virtual characters that can communicate naturally in a…

Cited by 4SourcePDFScholar
2024

ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling

IROS 2024

An environment acoustic model represents how sound is transformed by the physical characteristics of an indoor environment, for any given source/receiver location. Traditional methods for constructing acoustic models involve expensive and time-consuming collection of large quantities of acoustic dat

Cited by 2SourceScholar
2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2024

F3Loc: Fusion and Filtering for Floorplan Localization

CVPR 2024highlight

In this paper we propose an efficient data-driven solution to self-localization within a floorplan. Floorplan data is readily available long-term persistent and inherently robust to changes in the visual appearance. Our method does not require retraining per map and location or demand a large databa…

Cited by 7SourcePDFScholar
2024

HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness

NeurIPS 2024poster

We study the problem of precisely swapping objects in videos, with a focus on those interacted with by hands, given one user-provided reference object image. Despite the great advancements that diffusion models have made in video editing recently, these models often fall short in handling the intric…

Cited by 8SourcePDFScholar
2024

Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction

IROS 2024poster

Sim2real transfer has received increasing attention lately due to its success in transferring robotic policies learned in simulation to the real world. While significant progress has been made in transferring vision-based navigation policies, the current sim2real strategy for audio-visual navigation…

Cited by 4SourceScholar
2024

SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos

CVPR 2024poster

We propose a novel self-supervised embedding to learn how actions sound from narrated in-the-wild egocentric videos. Whereas existing methods rely on curated data with known audio-visual correspondence our multimodal contrastive-consensus coding (MC3) embedding reinforces the associations between au…

Cited by 8SourcePDFScholar
2023

Measuring Acoustics with Collaborative Multiple Agents

IJCAI 2023poster

As humans, we hear sound every second of our life. The sound we hear is often affected by the acoustics of the environment surrounding us. For example, a spacious hall leads to more reverberation. Room Impulse Responses (RIR) are commonly used to characterize environment acoustics as a function of t…

Cited by 3SourcePDFScholar
2023

Novel-View Acoustic Synthesis

CVPR 2023poster

We introduce the novel-view acoustic synthesis (NVAS) task: given the sight and sound observed at a source viewpoint, can we synthesize the sound of that scene from an unseen target viewpoint? We propose a neural rendering approach: Visually-Guided Acoustic Synthesis (ViGAS) network that learns to s…

2023

Overview of the L3DAS23 Challenge on Audio-Visual Extended Reality

ICASSP 2023accepted

The primary goal of the L3DAS23 Signal Processing Grand Challenge at ICASSP 2023 is to promote and support collaborative research on machine learning for 3D audio signal processing, with a specific emphasis on 3D speech enhancement and 3D Sound Event Localization and Detection in Extended Reality ap…

Cited by 0SourceScholar
2023

Replay: Multi-modal Multi-view Acted Videos for Casual Holography

ICCV 2023poster

We introduce Replay, a collection of multi-view, multi-modal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with several static cameras, as well as wearable action cameras, and recorded with a large array of microphones at different…

Cited by 7PDFcodeScholar
2023

SMUG Planner: A Safe Multi-Goal Planner for Mobile Robots in Challenging Environments

RA-L 2023

Robotic exploration or monitoring missions require mobile robots to autonomously and safely navigate between multiple target locations in potentially challenging environments. Currently, this type of multi-goal mission often relies on humans designing a set of actions for the robot to follow in the

Cited by 15SourcecodeScholar
2022

Few-Shot Audio-Visual Learning of Environment Acoustics

NeurIPS 2022accept

Room impulse response (RIR) functions capture how the surrounding physical environment transforms the sounds heard by a listener, with implications for various applications in AR, VR, and robotics. Whereas traditional methods to estimate RIRs assume dense geometry and/or sound measurements throughou…

Cited by 55SourcePDFScholar
2022

Sound Adversarial Audio-Visual Navigation

ICLR 2022poster

Audio-visual navigation task requires an agent to find a sound source in a realistic, unmapped 3D environment by utilizing egocentric audio-visual observations. Existing audio-visual navigation works assume a clean environment that solely contains the target sound, which, however, would not be suita…

2022

SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic Learning

NeurIPS 2022accept

We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from arbitrary microphone locations. Together with existing 3D vi…

Cited by 94SourcePDFScholar
2021

Learning to Set Waypoints for Audio-Visual Navigation

ICLR 2021poster

In audio-visual navigation, an agent intelligently travels through a complex, unmapped 3D environment using both sights and sounds to find a sound source (e.g., a phone ringing in another room). Existing models learn to act at a fixed granularity of agent motion and rely on simple recurrent aggregat…

2020

SoundSpaces: Audio-Visual Navigation in 3D Environments

ECCV 2020poster

Moving around in the world is naturally a multi-sensory experience, but today's embodied agents are deaf - restricted to solely their visual perception of the environment. We introduce audio-visual navigation for complex, acoustically and visually realistic 3D environments. By both seeing and hearin…

2020

VisualEchoes: Spatial Image Representation Learning through Echolocation

ECCV 2020poster

Several animal species (e.g., bats, dolphins, and whales) and even visually impaired humans have the remarkable ability to perform echolocation: a biological sonar used to perceive spatial layout and locate objects in the world. We explore the spatial cues contained in echoes and how they can benefi…

2019

Crowd-Robot Interaction: Crowd-Aware Robot Navigation With Attention-Based Deep Reinforcement Learning

ICRA 2019poster

Mobility in an effective and socially-compliant manner is an essential yet challenging task for robots operating in crowded spaces. Recent works have shown the power of deep reinforcement learning techniques to learn socially cooperative policies. However, their cooperation ability deteriorates as t…

Cited by 715SourceScholar