← Search

Ronghui Li

11 accepted papers

2025

A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions

ICCV 2025poster

Extracting physically plausible 3D human motion from videos is a critical task. Although existing simulation-based motion imitation methods can enhance the physical quality of daily motions estimated from monocular video capture, extending this capability to high-difficulty motions remains an open c…

2025

AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward

CVPR 2025poster

Recently, text-to-motion models open new possibilities for creating realistic human motion with greater efficiency and flexibility. However, aligning motion generation with event-level textual descriptions presents unique challenges due to the complex, nuanced relationship between textual prompts an…

Cited by 1SourcePDFScholar
2025

Harmonious Music-driven Group Choreography with Trajectory-Controllable Diffusion

AAAI 2025technical

Creating group choreography from music is crucial in cultural entertainment and virtual reality, with a focus on generating harmonious movements. Despite growing interest, recent approaches often struggle with two major challenges: multi-dancer collisions and single-dancer foot sliding. To address…

Cited by 0SourcePDFScholar
2025

Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling

ICCV 2025poster

Well-coordinated, music-aligned holistic dance enhances emotional expressiveness and audience engagement. However, generating such dances remains challenging due to the scarcity of holistic 3D dance datasets, the difficulty of achieving cross-modal alignment between music and dance, and the complexi…

2024

Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control

AAAI 2024technical

This study aims to improve the generation of 3D gestures by utilizing multimodal information from human speech. Previous studies have focused on incorporating additional modalities to enhance the quality of generated gestures. However, these methods perform poorly when certain modalities are missing…

Cited by 14SourcePDFScholar
2024

Cross-Modal Match for Language Conditioned 3D Object Grounding

AAAI 2024technical

Language conditioned 3D object grounding aims to find the object within the 3D scene mentioned by natural language descriptions, which mainly depends on the matching between visual and natural language. Considerable improvement in grounding performance is achieved by improving the multimodal fusion…

Cited by 9SourcePDFScholar
2024

Exploring Multi-Modal Control in Music-Driven Dance Generation

ICASSP 2024accepted

Existing music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating high-quality dance movements and supporting multi-modal cont…

Cited by 0SourceScholar
2024

Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance Primitives

CVPR 2024poster

We propose Lodge a network capable of generating extremely long dance sequences conditioned on given music. We design Lodge as a two-stage coarse to fine diffusion architecture and propose the characteristic dance primitives that possess significant expressiveness as intermediate representations bet…

2024

MambaTalk: Efficient Holistic Gesture Synthesis with Selective State Space Models

NeurIPS 2024poster

Gesture synthesis is a vital realm of human-computer interaction, with wide-ranging applications across various fields like film, robotics, and virtual reality. Recent advancements have utilized the diffusion model to improve gesture synthesis. However, the high computational complexity of these t…

2024

Text2Avatar: Text to 3d Human Avatar Generation with Codebook-Driven Body Controllable Attribute

ICASSP 2024accepted

Generating 3D human models directly from text helps reduce the cost and time of character modeling. However, achieving multi-attribute controllable and realistic 3D human avatar generation is still challenging due to feature coupling and the scarcity of realistic 3D human avatar datasets. To address…

Cited by 0SourceScholar
2023

FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance Generation

ICCV 2023poster

Generating full-body and multi-genre dance sequences from given music is a challenging task, due to the limitations of existing datasets and the inherent complexity of the fine-grained hand motion and dance genres. To address these problems, we propose FineDance, which contains 14.6 hours of music-…

Cited by 58PDFcodeScholar