← Search

Shaoli Huang

29 accepted papers

2026

From 2D Alignment to 3D Plausibility: Unifying Heterogeneous 2D Priors and Penetration-Free Diffusion for Occlusion-Robust Two-Hand Reconstruction

CVPR 2026

Two-hand reconstruction from monocular images is hampered by complex poses and severe occlusions, which often cause interaction misalignment and two-hand penetration. We address this by decoupling the problem into 2D structural alignment and 3D spatial interaction alignment, each handled by a tailor

Cited by 0SourcecodeScholar
2026

Human-Object Interaction via Automatically Designed VLM-Guided Motion Policy

ICLR 2026poster

Human-object interaction (HOI) synthesis is crucial for applications in animation, simulation, and robotics. However, existing approaches either rely on expensive motion capture data or require manual reward engineering, limiting their scalability and generalizability. In this work, we introduce the…

Cited by 0SourcecodeScholar
2025

Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection

ICCV 2025poster

Open vocabulary Human-Object Interaction (HOI) detection is a challenging task that detects all <human, verb, object> triplets of interest in an image, even those that are not pre-defined in the training set. Existing approaches typically rely on output features generated by large Vision-Language Mo…

2025

DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time Gesture Generation from Speech

AAAI 2025technical

Diffusion models have demonstrated remarkable synthesis quality and diversity in generating co-speech gestures. However, the computationally intensive sampling steps associated with diffusion models hinder their practicality in real-world applications. Hence, we present DIDiffGes, for a Decoupled…

Cited by 0SourcePDFScholar
2025

Democratizing High-Fidelity Co-Speech Gesture Video Generation

ICCV 2025poster

Co-speech gesture video generation aims to synthesize realistic, audio-aligned videos of speakers, complete with synchronized facial expressions and body gestures. This task presents challenges due to the significant one-to-many mapping between audio and visual content, further complicated by the sc…

2025

Guiding Human-Object Interactions with Rich Geometry and Relations

CVPR 2025poster

Human-object interaction (HOI) synthesis is crucial for creating immersive and realistic experiences for applications such as virtual reality. Existing methods often rely on simplified object representations, such as the object's centroid or the nearest point to a human, to achieve physically plausi…

Cited by 1SourcePDFScholar
2024

Exploring the Feature Extraction and Relation Modeling For Light-Weight Transformer Tracking

ECCV 2024oral

"Recent advancements in transformer-based lightweight object tracking have set new standards across various benchmarks due to their efficiency and effectiveness. Despite these achievements, most current trackers rely heavily on pre-existing object detection architectures without optimizing the backb…

2024

FreeTalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness

ICASSP 2024accepted

Current talking avatars mostly generate co-speech gestures based on audio and text of the utterance, without considering the non-speaking motion of the speaker. Furthermore, previous works on co-speech gesture generation have designed network structures based on individual gesture datasets, which re…

Cited by 0SourceScholar
2024

HuTuMotion: Human-Tuned Navigation of Latent Motion Diffusion Models with Minimal Feedback

AAAI 2024technical

We introduce HuTuMotion, an innovative approach for generating natural human motions that navigates latent motion diffusion models by leveraging few-shot human feedback. Unlike existing approaches that sample latent variables from a standard normal prior distribution, our method adapts the prior dis…

Cited by 6SourcePDFScholar
2024

MambaTree: Tree Topology is All You Need in State Space Model

NeurIPS 2024spotlight

The state space models, employing recursively propagated features, demonstrate strong representation capabilities comparable to Transformer models and superior efficiency. However, constrained by the inherent geometric constraints of sequences, it still falls short in modeling long-range dependencie…

Cited by 1SourcePDFScholar
2024

Programmable Motion Generation for Open-Set Motion Control Tasks

CVPR 2024highlight

Character animation in real-world scenarios necessitates a variety of constraints such as trajectories key-frames interactions etc. Existing methodologies typically treat single or a finite set of these constraint(s) as separate control tasks. These methods are often specialized and the tasks they a…

Cited by 5SourcePDFScholar
2024

TapMo: Shape-aware Motion Generation of Skeleton-free Characters

ICLR 2024poster

Previous motion generation methods are limited to the pre-rigged 3D human model, hindering their applications in the animation of various non-rigged characters. In this work, we present TapMo, a Text-driven Animation PIpeline for synthesizing Motion in a broad spectrum of skeleton-free 3D characters…

Cited by 11SourcePDFScholar
2023

ACR: Attention Collaboration-Based Regressor for Arbitrary Two-Hand Reconstruction

CVPR 2023poster

Reconstructing two hands from monocular RGB images is challenging due to frequent occlusion and mutual confusion. Existing methods mainly learn an entangled representation to encode two interacting hands, which are incredibly fragile to impaired interaction, such as truncated hands, separate hands,…

2023

Harmonious Feature Learning for Interactive Hand-Object Pose Estimation

CVPR 2023poster

Joint hand and object pose estimation from a single image is extremely challenging as serious occlusion often occurs when the hand and object interact. Existing approaches typically first extract coarse hand and object features from a single backbone, then further enhance them with reference to each…

2023

Learning Anchor Transformations for 3D Garment Animation

CVPR 2023poster

This paper proposes an anchor-based deformation model, namely AnchorDEF, to predict 3D garment animation from a body motion sequence. It deforms a garment mesh template by a mixture of rigid transformations with extra nonlinear displacements. A set of anchors around the mesh surface is introduced to…

Cited by 13SourcePDFScholar
2023

LivelySpeaker: Towards Semantic-Aware Co-Speech Gesture Generation

ICCV 2023poster

Gestures are non-verbal but important behaviors accompanying people's speech. While previous methods are able to generate speech rhythm-synchronized gestures, the semantic context of the speech is generally lacking in the gesticulations. Although semantic gestures do not occur very regularly in huma…

Cited by 26PDFcodeScholar
2023

LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior Understanding

ICCV 2023poster

Understanding and analyzing animal behavior is increasingly essential to protect endangered animal species. However, the application of advanced computer vision techniques in this regard is minimal, which boils down to lacking large and diverse datasets for training deep models. To break the deadloc…

Cited by 18PDFcodeScholar
2023

Master: Meta Style Transformer for Controllable Zero-Shot and Few-Shot Artistic Style Transfer

CVPR 2023poster

Transformer-based models achieve favorable performance in artistic style transfer recently thanks to its global receptive field and powerful multi-head/layer attention operations. Nevertheless, the over-paramerized multi-layer structure increases parameters significantly and thus presents a heavy bu…

Cited by 20SourcePDFScholar
2023

Skinned Motion Retargeting With Residual Perception of Motion Semantics & Geometry

CVPR 2023poster

A good motion retargeting cannot be reached without reasonable consideration of source-target differences on both the skeleton and shape geometry levels. In this work, we propose a novel Residual RETargeting network (R2ET) structure, which relies on two neural modification modules, to adjust the sou…

2022

Geometric Structure Preserving Warp for Natural Image Stitching

CVPR 2022poster

Preserving geometric structures in the scene plays a vital role in image stitching. However, most of the existing methods ignore the large-scale layouts reflected by straight lines or curves, decreasing overall stitching quality. To address this issue, this work presents a structure-preserving stitc…

Cited by 41PDFcodeScholar
2022

Towards Hard-Positive Query Mining for DETR-Based Human-Object Interaction Detection

ECCV 2022poster

"Human-Object Interaction (HOI) detection is a core task for high-level image understanding. Recently, Detection Transformer (DETR)-based HOI detectors have become popular due to their superior performance and efficient structure. However, these approaches typically adopt fixed HOI queries for all t…

2021

SnapMix: Semantically Proportional Mixing for Augmenting Fine-grained Data

AAAI 2021technical

Data mixing augmentation has proved effective in training deep models. Recent methods mix labels mainly according to the mixture proportion of image pixels. Due to the major discriminative information of a fine-grained image usually resides in subtle regions, these methods tend to introduce heavy la…

2021

Stochastic Partial Swap: Enhanced Model Generalization and Interpretability for Fine-Grained Recognition

ICCV 2021poster

Learning mid-level representation for fine-grained recognition is easily dominated by a limited number of highly discriminative patterns, degrading its robustness and generalization capability. To this end, we propose a novel Stochastic Partial Swap (SPS) scheme to address this issue. Our method per…

Cited by 67PDFcodeScholar
2019

Learning a Mixture of Granularity-Specific Experts for Fine-Grained Categorization

ICCV 2019poster

We aim to divide the problem space of fine-grained recognition into some specific regions. To achieve this, we develop a unified framework based on a mixture of experts. Due to limited data available for the fine-grained recognition problem, it is not feasible to learn diverse experts by using a dat…

Cited by 221PDFScholar
2019

Not All Parts Are Created Equal: 3D Pose Estimation by Modeling Bi-Directional Dependencies of Body Parts

ICCV 2019oral

Not all the human body parts have the same degree of freedom (DOF) due to the physiological structure. For example, the limbs may move more flexibly and freely than the torso does. Most of the existing 3D pose estimation methods, despite the very promising results achieved, treat the body joints equ…

Cited by 74PDFScholar