← Search

Jiaben Chen

11 accepted papers

2025

RapVerse: Coherent Vocals and Whole-Body Motion Generation from Text

ICCV 2025poster

In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these two modalities in isolation. To facilitate this, we first collect the RapVerse dat…

Cited by 0SourcePDFScholar
2025

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

NeurIPS 2025poster

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips totaling over 500 hours of high-quality 1080P human speech videos w…

Cited by 0SourceScholar
2025

UniMuMo: Unified Text, Music, and Motion Generation

AAAI 2025technical

We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired music and motion data based on rhythmic patterns to leverage…

2024

Diffusion-Generated Pseudo-Observations for High-Quality Sparse-View Reconstruction

ECCV 2024poster

"Novel view synthesis via Neural Radiance Fields (NeRFs) or 3D Gaussian Splatting (3DGS) typically necessitates dense observations with hundreds of input images to circumvent artifacts. We introduce Deceptive-NeRF/3DGS1 to enhance sparse-view reconstruction with only a limited set of input images, b…

Cited by 0SourcePDFScholar
2024

RoboDreamer: Learning Compositional World Models for Robot Imagination

ICML 2024poster

Text-to-video models have demonstrated substantial potential in robotic decision-making, enabling the imagination of realistic plans of future actions as well as accurate environment simulation. However, one major issue in such models is generalization -- models are limited to synthesizing videos su…

Cited by 23SourcePDFScholar
2023

Revisiting Event-Based Video Frame Interpolation

IROS 2023poster

Dynamic vision sensors or event cameras provide rich complementary information for video frame interpolation. Existing state-of-the-art methods follow the paradigm of combining both synthesis-based and warping networks. However, few of those methods fully respect the intrinsic characteristics of eve…

Cited by 4SourceScholar
2023

iQuery: Instruments As Queries for Audio-Visual Sound Separation

CVPR 2023poster

Current audio-visual separation methods share a standard architecture design where an audio encoder-decoder network is fused with visual encoding features at the encoder bottleneck. This design confounds the learning of multi-modal feature encoding with robust sound decoding for audio separation. To…

2022

AutoVideo: An Automated Video Action Recognition System

IJCAI 2022poster

Action recognition is an important task for video understanding with broad applications. However, developing an effective action recognition solution often requires extensive engineering efforts in building and testing different combinations of the modules and their hyperparameters. In this demo, we…

2022

DEVO: Depth-Event Camera Visual Odometry in Challenging Conditions

ICRA 2022poster

We present a novel real-time visual odometry framework for a stereo setup of a depth and high-resolution event camera. Our framework balances accuracy and robustness against computational efficiency towards strong performance in challenging scenarios. We extend conventional edge-based semi-dense vis…

Cited by 65SourceScholar
2022

Unsupervised Multi-View Object Segmentation Using Radiance Field Propagation

NeurIPS 2022accept

We present radiance field propagation (RFP), a novel approach to segmenting objects in 3D during reconstruction given only unlabeled multi-view images of a scene. RFP is derived from emerging neural radiance field-based techniques, which jointly encodes semantics with appearance and geometry. The co…

Cited by 30SourcePDFScholar
2022

VECtor: A Versatile Event-Centric Benchmark for Multi-Sensor SLAM

RA-L 2022

Event cameras have recently gained in popularity as they hold strong potential to complement regular cameras in situations of high dynamics or challenging illumination. An important problem that may benefit from the addition of an event camera is given by Simultaneous Localization And Mapping (SLAM)

Cited by 103SourceScholar