← Search

Zhanbo Huang

5 accepted papers

2026

Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait Recognition

CVPR 2026

Existing set-based gait recognition methods achieve remarkable performance by capturing global semantic context.However, their order-invariant nature prevents them from modeling the fine-grained kinematic patterns that unfold over time.To unify the global and process-level representations, we propos

Cited by 0SourceScholar
2025

BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision Models

NeurIPS 2025poster

Large vision models (LVM) based gait recognition has achieved impressive performance. However, existing LVM-based approaches may overemphasize gait priors while neglecting the intrinsic value of LVM itself, particularly the rich, distinct representations across its multi-layers. To adequately unloc…

Cited by 0SourcecodeScholar
2025

H-MoRe: Learning Human-centric Motion Representation for Action Analysis

CVPR 2025highlight

In this paper, we propose H-MoRe, a novel pipeline for learning precise human-centric motion representation. Our approach dynamically preserves relevant human motion while filtering out background movement. Notably, unlike previous methods relying on fully supervised learning from synthetic data, H-…

2022

ReCoNet: Recurrent Correction Network for Fast and Efficient Multi-Modality Image Fusion

ECCV 2022poster

"Recent advances in deep networks have gained great attention in infrared and visible image fusion (IVIF). Nevertheless, most existing methods are incapable of dealing with slight misalignment on source images and suffer from high computational and spatial expenses. This paper tackles these two crit…

2022

Target-Aware Dual Adversarial Learning and a Multi-Scenario Multi-Modality Benchmark To Fuse Infrared and Visible for Object Detection

CVPR 2022oral

This study addresses the issue of fusing infrared and visible images that appear differently for object detection. Aiming at generating an image of high visual quality, previous approaches discover commons underlying the two modalities and fuse upon the common space either by iterative optimization…

Cited by 733PDFcodeScholar