← Search

Aimin Hao

15 accepted papers

2026

DECON: Reconstruction of Clothed-Geometric Multiple Humans from a Single Image via Geometry-Guided Decoupling

AAAI 2026technical

3D multi-human reconstruction from single images holds significant potential for advancing AR/VR applications. While remarkable progress has been made in single-human reconstruction, existing methods face challenges when reconstructing multiple humans. These challenges include: (1) severe inter-occl

Cited by 0SourcePDFScholar
2026

IntentMotion: Learning Intent-Aware Human Motion from Language in 3D Scenes

AAAI 2026technical

Generating human motion in complex 3D scenes from text is a challenging task with broad applications. However, existing methods often overlook realistic physical contact, resulting in visually plausible but physically unrealistic motion, e.g., penetration. To alleviate this, we propose IntentMotion,

Cited by 0SourcePDFScholar
2026

MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion Generation

CVPR 2026

Diffusion-based motion generation has advanced rapidly, but current methods still struggle with long-horizon consistency, style control, and multi-condition guidance. A major reason is the fused-conditioning design, where semantic, stylistic, and temporal signals share a single pathway, causing inte

Cited by 0SourceScholar
2026

TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition

CVPR 2026

Understanding complex surgical scenes requires recognizing multiple interdependent entities--such as instruments, actions, and targets--and maintaining their relational consistency across time. Existing surgical triplet recognition methods struggle to jointly model intra-frame label dependencies and

Cited by 0SourceScholar
2025

3D Dental Model Segmentation with Geometrical Boundary Preserving

CVPR 2025poster

3D intraoral scan mesh is widely used in digital dentistry diagnosis, segmenting 3D intraoral scan mesh is a critical preliminary task. Numerous approaches have been devised for precise tooth segmentation. Currently, the deep learning-based methods are capable of the high accuracy segmentation of c…

2025

CtrlAvatar: Controllable Avatars Generation via Disentangled Invertible Networks

AAAI 2025technical

As virtual experiences grow in popularity, the demand for realistic, personalized, and animatable human avatars increases. Traditional methods, relying on fixed templates, often produce costly avatars that lack expressiveness and realism. To overcome these challenges, we introduce Controllable Avata…

2024

Arbitrary Motion Style Transfer with Multi-condition Motion Latent Diffusion Model

CVPR 2024poster

Computer animation's quest to bridge content and style has historically been a challenging venture with previous efforts often leaning toward one at the expense of the other. This paper tackles the inherent challenge of content-style duality ensuring a harmonious fusion where the core narrative of t…

2024

FaceCom: Towards High-fidelity 3D Facial Shape Completion via Optimization and Inpainting Guidance

CVPR 2024poster

We propose FaceCom a method for 3D facial shape completion which delivers high-fidelity results for incomplete facial inputs of arbitrary forms. Unlike end-to-end shape completion methods based on point clouds or voxels our approach relies on a mesh-based generative network that is easy to optimize…

2024

HOIAnimator: Generating Text-prompt Human-object Animations using Novel Perceptive Diffusion Models

CVPR 2024poster

To date the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehensive physics-centric…

Cited by 11SourcePDFScholar
2024

Weakly Supervised Multimodal Affordance Grounding for Egocentric Images

AAAI 2024technical

To enhance the interaction between intelligent systems and the environment, locating the affordance regions of objects is crucial. These regions correspond to specific areas that provide distinct functionalities. Humans often acquire the ability to identify these regions through action demonstration…

2023

Pixel Is All You Need: Adversarial Trajectory-Ensemble Active Learning for Salient Object Detection

AAAI 2023technical

Although weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by…

Cited by 10SourcePDFScholar
2023

Sequential Texts Driven Cohesive Motions Synthesis with Natural Transitions

ICCV 2023poster

The intelligent synthesis/generation of daily-life motion sequences is fundamental and urgently needed for many VR/metaverse-related applications. However, existing approaches commonly focus on monotonic motion generation (e.g., walking, jumping, etc.) based on single instruction-like text, which is…

Cited by 15PDFcodeScholar
2021

From Semantic Categories to Fixations: A Novel Weakly-Supervised Visual-Auditory Saliency Detection Approach

CVPR 2021poster

Thanks to the rapid advances in the deep learning techniques and the wide availability of large-scale training sets, the performances of video saliency detection models have been improving steadily and significantly. However, the deep learning based visual-audio fixation prediction is still in its i…

Cited by 47PDFcodeScholar