← Search

Man Zhang

15 accepted papers

2026

A3fford-HOI: Anatomy-Aligned Affordance Disentanglement for Fine-grained and Generalizable Hand Object Interaction

IJCAI 2026

Hand Object Interaction (HOI) generation provides an efficient solution for virtual reality simulation and embodied AI deployment.Recent studies have explored instruction-driven HOI synthesis, yet they overlooked the fine-grained interactive contact and struggled with robust generalization in data-s

Cited by 0Scholar
2026

Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models

ICLR 2026poster

Text-to-image (T2I) models have achieved remarkable success in generating high-fidelity images, but they often fail in handling complex spatial relationships, e.g., spatial perception, reasoning, or interaction. These critical aspects are largely overlooked by current benchmarks due to their short o…

Cited by 0SourcecodeScholar
2026

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction

AAAI 2026technical

Generating responsive listener head dynamics with nuanced emotions and expressive reactions is crucial for dialogue modeling in various virtual avatar animations. Previous studies mainly focus on the direct short-term production of listener behavior. They overlook the fine-grained control over motio

Cited by 0SourcePDFScholar
2025

DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions

ICCV 2025poster

Generating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis, they fail to recognize that enabling users to edit dance movements is far more practical in real-world choreography scen…

2025

Gait-X: Exploring X modality for Generalized Gait Recognition

ICCV 2025poster

Modality exploration has been repeatedly mentioned in gait recognition, evolving from silhouette to parsing, mesh, point clouds, etc. These latest modalities agree that silhouette is less affected by background and clothing noises, but argue it loses too much valuable discriminative information. The…

Cited by 0SourcePDFScholar
2025

OpenAnimals: Revisiting Person Re-Identification for Animals Towards Better Generalization

ICCV 2025poster

This paper addresses the challenge of animal re-identification, an emerging field that shares similarities with person re-identification but presents unique complexities due to the diverse species, environments and poses. To facilitate research in this domain, we introduce OpenAnimals, a flexible an…

2024

DefFusion: Deformable Multimodal Representation Fusion for 3D Semantic Segmentation

ICRA 2024poster

The complementarity between camera and LiDAR data makes fusion methods a promising approach to improve 3D semantic segmentation performance. Recent transformer-based methods have also demonstrated superiority in segmentation. However, multimodal solutions incorporating transformers are underexplored…

Cited by 8SourceScholar
2024

HardMo: A Large-Scale Hardcase Dataset for Motion Capture

CVPR 2024poster

Recent years have witnessed rapid progress in monocular human mesh recovery. Despite their impressive performance on public benchmarks existing methods are vulnerable to unusual poses which prevents them from deploying to challenging scenarios such as dance and martial arts. This issue is mainly att…

Cited by 1SourcePDFScholar
2024

ISP-Teacher:Image Signal Process with Disentanglement Regularization for Unsupervised Domain Adaptive Dark Object Detection

AAAI 2024technical

Object detection in dark conditions has always been a great challenge due to the complex formation process of low-light images. Currently, the mainstream methods usually adopt domain adaptation with Teacher-Student architecture to solve the dark object detection problem, and they imitate the dark co…

2024

Probabilistic Contrastive Learning for Domain Adaptation

IJCAI 2024poster

Contrastive learning has shown impressive success in enhancing feature discriminability for various visual tasks in a self-supervised manner, but the standard contrastive paradigm (features+l2 normalization) has limited benefits when applied in domain adaptation. We find that this is mainly because…

2024

QAGait: Revisit Gait Recognition from a Quality Perspective

AAAI 2024technical

Gait recognition is a promising biometric method that aims to identify pedestrians from their unique walking patterns. Silhouette modality, renowned for its easy acquisition, simple structure, sparse representation, and convenient modeling, has been widely employed in controlled in-the-lab research.…

2024

RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

ACL 2024findings

The advent of Large Language Models (LLMs) has paved the way for complex tasks such as role-playing, which enhances user interactions by enabling models to imitate various characters. However, the closed-source nature of state-of-the-art LLMs and their general-purpose training limit role-playing opt…

2024

Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic Segmentation

AAAI 2024technical

Recently, CLIP has found practical utility in the domain of pixel-level zero-shot segmentation tasks. The present landscape features two-stage methodologies beset by issues such as intricate pipelines and elevated computational costs. While current one-stage approaches alleviate these concerns and…

2023

DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action Localization

ICCV 2023poster

Weakly-supervised temporal action localization (WTAL) is a practical yet challenging task. Due to large-scale datasets, most existing methods use a network pretrained in other datasets to extract features, which are not suitable enough for WTAL. To address this problem, researchers design several mo…

Cited by 18PDFcodeScholar