← Search

Qian Bao

7 accepted papers

2026

Multi-level Causal LLM-based Text-to-Motion Generation with Human Alignment

CVPR 2026

Although progress has been made in LLM-based text-driven motion generation, it still has the limitations of generating fine-grained and semantically consistent motions. These limitations stem from: 1) fine-grained motion quantization errors; 2) mismatches between causal reasoning language and non-ca

Cited by 0SourceScholar
2023

TRACE: 5D Temporal Regression of Avatars With Dynamic Cameras in 3D Environments

CVPR 2023poster

Although the estimation of 3D human pose and shape (HPS) is rapidly progressing, current methods still cannot reliably estimate moving humans in global coordinates, which is critical for many applications. This is particularly challenging when the camera is also moving, entangling human and camera m…

2022

Genre-Conditioned Long-Term 3D Dance Generation Driven by Music

ICASSP 2022accepted

Dancing to music is an artistic behavior of humans, however, letting machines generate dances from music is still challenging. Most existing works have been made progress in tackling the problem of motion prediction conditioned by music, yet they rarely consider the importance of the musical genre.…

Cited by 0SourceScholar
2022

Learning Monocular Mesh Recovery of Multiple Body Parts Via Synthesis

ICASSP 2022accepted

In this paper, we focus on simultaneously recovering the 3D mesh of multiple body parts from a single RGB image. One of the main challenges is that available datasets with full-body 3D annotations are very limited. This results in poor generalization ability of existing learning-based methods. Exist…

Cited by 0SourceScholar
2022

Putting People in Their Place: Monocular Regression of 3D People in Depth

CVPR 2022poster

Given an image with multiple people, our goal is to directly regress the pose and shape of all the people as well as their relative depth. Inferring the depth of a person in an image, however, is fundamentally ambiguous without knowing their height. This is particularly problematic when the scene co…

Cited by 180PDFcodeScholar
2021

Monocular, One-Stage, Regression of Multiple 3D People

ICCV 2021poster

This paper focuses on the regression of multiple 3D people from a single RGB image. Existing approaches predominantly follow a multi-stage pipeline that first detects people in bounding boxes and then independently regresses their 3D body meshes. In contrast, we propose to Regress all meshes in a On…

Cited by 327PDFcodeScholar
2021

Neural Architecture Search for Joint Human Parsing and Pose Estimation

ICCV 2021poster

Human parsing and pose estimation are crucial for the understanding of human behaviors. Since these tasks are closely related, employing one unified model to perform two tasks simultaneously allows them to benefit from each other. However, since human parsing is a pixel-wise classification process w…

Cited by 26PDFcodeScholar