← Search

Xiaohu Huang

7 accepted papers

2026

OmniShow: Orchestrating Multimodal Conditions for Human-Object Interaction Video Generation

ICML 2026poster

In this work, we study **Human-Object Interaction Video Generation (HOIVG)**, which aims to synthesize high-quality HOI videos via text, reference image, audio, and pose conditions. To address the challenges of harmonious multimodal injection and heterogeneous data utility, we present **OmniShow**, …

Cited by 0SourceScholar
2025

Change3D: Revisiting Change Detection and Captioning from A Video Modeling Perspective

CVPR 2025highlight

In this paper, we present Change3D, a framework that reconceptualizes the change detection and captioning tasks through video modeling. Recent methods have achieved remarkable success by regarding each pair of bi-temporal images as separate frames. They employ a shared-weight image encoder to extrac…

2024

FROSTER: Frozen CLIP is A Strong Teacher for Open-Vocabulary Action Recognition

ICLR 2024poster

In this paper, we introduce \textbf{FROSTER}, an effective framework for open-vocabulary action recognition. The CLIP model has achieved remarkable success in a range of image-based tasks, benefiting from its strong generalization capability stemming from pretaining on massive image-text pairs. Howe…

2023

Graph Contrastive Learning for Skeleton-based Action Recognition

ICLR 2023poster

In the field of skeleton-based action recognition, current top-performing graph convolutional networks (GCNs) exploit intra-sequence context to construct adaptive graphs for feature aggregation. However, we argue that such context is still $\textit{local}$ since the rich cross-sequence relations hav…

2021

Context-Sensitive Temporal Feature Learning for Gait Recognition

ICCV 2021poster

Although gait recognition has drawn increasing research attention recently, it remains challenging to learn discriminative temporal representation since the silhouette differences are quite subtle in spatial domain. Inspired by the observation that humans can distinguish gaits of different subjects…

Cited by 169PDFcodeScholar