← Search

Ju Dai

9 accepted papers

2026

VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network

CVPR 2026

The rapid advances in deep learning have significantly enhanced the accuracy of multimodal 3D human pose estimation (HPE). However, the state-of-the-art (SOTA) HPE pipelines still rely on Transformers, whose quadratic complexity makes real-time processing for long sequences impractical. Mamba addres

Cited by 0SourcecodeScholar
2025

AU-Blendshape for Fine-grained Stylized 3D Facial Expression Manipulation

ICCV 2025poster

While 3D facial animation has made impressive progress, challenges still exist in realizing fine-grained stylized 3D facial expression manipulation due to the lack of appropriate datasets. In this paper, we introduce the AUBlendSet, a 3D facial dataset based on AU-Blendshape representation for fine-…

2025

Chat-Driven 3D Human Pose and Shape Editing with Large Language Models

ICASSP 2025accepted

Generating and creating humanoid 3D models has received increasing attention recently due to its fundamental support for many high-level 3D applications. Although automatic 3D pose and shape reconstruction methods have achieved promising results, there are still some failure cases due to self-occlus…

Cited by 0SourceScholar
2025

Hierarchical Proxy Learning for Cloth-Changing Person Re-Identification

ICASSP 2025accepted

Cloth-Changing person Re-Identification (CC-ReID) depends significantly on learning discriminative features under the cloth-changing scenario. It is quite challenging due to the large intra-person variance and small inter-person variance caused by clothes changing. To address these issues, in this w…

Cited by 0SourceScholar
2025

Position-Aware Guided Point Cloud Completion with CLIP Model

AAAI 2025technical

Point cloud completion aims to recover partial geometric and topological shapes caused by equipment defects or limited viewpoints. Current methods either solely rely on the 3D coordinates of the point cloud to complete it or incorporate additional images with well-calibrated intrinsic parameters to…

Cited by 0SourcePDFScholar
2025

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation

CVPR 2025poster

In 3D speech-driven facial animation generation, existing methods commonly employ pre-trained self-supervised audio models as encoders. However, due to the prevalence of phonetically similar syllables with distinct lip shapes in language, these near-homophone syllables tend to exhibit significant co…

2024

AHRNET: Attention and Heatmap-Based Regressor for Hand Pose Estimation and Mesh Recovery

ICASSP 2024accepted

Estimating 3D hand pose and recovering the full hand surface mesh from a single RGB image is a challenging task due to self-occlusions, viewpoint changes, and the complexity of hand articulations. In this paper, we propose a novel framework that combines an attention mechanism with heatmap regressio…

Cited by 0SourceScholar
2018

A Bi-Directional Message Passing Model for Salient Object Detection

CVPR 2018poster

Recent progress on salient object detection is beneficial from Fully Convolutional Neural Network (FCN). The saliency cues contained in multi-level convolutional features are complementary for detecting salient objects. How to integrate multi-level features becomes an open problem in saliency detect…

Cited by 579SourcePDFScholar