← Search

Peihao Chen

21 accepted papers

2026

NaVLA$^2$: A Vision-Language-Audio-Action Model for Multimodal Instruction Navigation

AAAI 2026technical

Embodied navigation is a fundamental capability for intelligent agents, yet remains challenging in partially observable environments where navigation instructions can be difficult to interpret. However, existing tasks only provide unimodal instructions, which are ambiguous in complex multimodal envi

Cited by 0SourcePDFScholar
2025

3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning

CVPR 2025poster

Constructing compact and informative 3D scene representations is essential for effective embodied exploration and reasoning, especially in complex environments over extended periods. Existing representations, such as object-centric 3D scene graphs, oversimplify spatial relationships by modeling scen…

Cited by 1SourcePDFScholar
2025

Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic Guidance

IJCAI 2025

Open-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal assistants. The advent of large language models (LLMs) has greatly advanced this field by improving context understanding a

2025

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

CVPR 2025poster

Research on 3D Vision-Language Models (3D-VLMs) is gaining increasing attention, which is crucial for developing embodied AI within 3D scenes, such as visual navigation and embodied question answering. Due to the high density of visual features, especially in large 3D scenes, accurately locating tas…

2025

Learning 3D Persistent Embodied World Models

NeurIPS 2025poster

The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing work has explored how to construct such world models using vid…

Cited by 0SourceScholar
2024

3D-VLA: A 3D Vision-Language-Action Generative World Model

ICML 2024poster

Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception to action, neglecting the vast dynamics of the world and the relations between a…

Cited by 77SourcePDFScholar
2024

CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

ICLR 2024poster

A remarkable ability of human beings resides in compositional reasoning, i.e., the capacity to make "infinite use of finite means". However, current large vision-language foundation models (VLMs) fall short of such compositional abilities due to their ``bag-of-words" behaviors and inability to cons…

Cited by 16SourcePDFScholar
2024

FlexAttention for Efficient High-Resolution Vision-Language Models

ECCV 2024poster

"Current high-resolution vision-language models encode images as high-resolution image tokens and exhaustively take all these tokens to compute attention, which significantly increases the computational cost. To address this problem, we propose , a flexible attention mechanism for efficient high-res…

Cited by 13SourcePDFScholar
2024

MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World

CVPR 2024poster

Human beings possess the capability to multiply a melange of multisensory cues while actively exploring and interacting with the 3D world. Current multi-modal large language models however passively absorb sensory data as inputs lacking the capacity to actively interact with the objects in the 3D en…

Cited by 34SourcePDFScholar
2024

RILA: Reflective and Imaginative Language Agent for Zero-Shot Semantic Audio-Visual Navigation

CVPR 2024poster

We leverage Large Language Models (LLM) for zeroshot Semantic Audio Visual Navigation (SAVN). Existing methods utilize extensive training demonstrations for reinforcement learning yet achieve relatively low success rates and lack generalizability. The intermittent nature of auditory signals further…

Cited by 7SourcePDFScholar
2023

3D-LLM: Injecting the 3D World into Large Language Models

NeurIPS 2023spotlight

Large language models (LLMs) and Vision-Language Models (VLMs) have been proved to excel at multiple tasks, such as commonsense reasoning. Powerful as these models can be, they are not grounded in the 3D physical world, which involves richer concepts such as spatial relationships, affordances, physi…

Cited by 328SourcePDFScholar
2023

FGPrompt: Fine-grained Goal Prompting for Image-goal Navigation

NeurIPS 2023poster

Learning to navigate to an image-specified goal is an important but challenging task for autonomous systems like household robots. The agent is required to well understand and reason the location of the navigation goal from a picture shot in the goal position. Existing methods try to solve this prob…

Cited by 13SourcePDFScholar
2023

Learning Vision-and-Language Navigation from YouTube Videos

ICCV 2023poster

Vision-and-language navigation (VLN) requires an embodied agent to navigate in realistic 3D environments using natural language instructions. Existing VLN methods suffer from training on small-scale environments or unreasonable path-instruction datasets, limiting the generalization to unseen environ…

Cited by 31PDFcodeScholar
2023

Masked Motion Encoding for Self-Supervised Video Representation Learning

CVPR 2023poster

How to learn discriminative video representation from unlabeled videos is challenging but crucial for video analysis. The latest attempts seek to learn a representation model by predicting the appearance contents in the masked regions. However, simply masking and recovering appearance contents may n…

2022

Learning Active Camera for Multi-Object Navigation

NeurIPS 2022accept

Getting robots to navigate to multiple objects autonomously is essential yet difficult in robot applications. One of the key challenges is how to explore environments efficiently with camera sensors only. Existing navigation methods mainly focus on fixed cameras and few attempts have been made to na…

Cited by 27SourcePDFScholar
2022

Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation

NeurIPS 2022accept

We address a practical yet challenging problem of training robot agents to navigate in an environment following a path described by some language instructions. The instructions often contain descriptions of objects in the environment. To achieve accurate and efficient navigation, it is critical to b…

2021

RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning

AAAI 2021technical

We study unsupervised video representation learning that seeks to learn both motion and appearance features from unlabeled video only, which can be reused for downstream tasks such as action recognition. This task, however, is extremely challenging due to 1) the highly complex spatial-temporal infor…

2020

Foley Music: Learning to Generate Music from Videos

ECCV 2020poster

In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments. We first identify two key intermediate representations for a successful video to music generator: body keypoints from videos and MIDI events from au…

Cited by 168SourcePDFScholar
2019

Self-Supervised Moving Vehicle Tracking With Stereo Sound

ICCV 2019poster

Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audiovisual data to learn to localize objects (moving vehicles) in a visual re…

Cited by 174PDFScholar