← Search

Chenghao Xu

17 accepted papers

2026

Beyond Global Alignment: Fine-Grained Motion-Language Retrieval via Pyramidal Shapley-Taylor Learning

ICML 2026spotlight

As a foundational task in human-centric cross-modal intelligence, motion-language retrieval aims to bridge the semantic gap between natural language and human motion, enabling intuitive motion analysis, yet existing approaches predominantly focus on aligning entire motion sequences with global textu…

Cited by 0SourceScholar
2026

Channel-masked Asymmetric Distribution Matching for Cross-Domain Generalized Dataset Distillation

AAAI 2026technical

Dataset distillation has achieved remarkable progress as an effective approach for data compression. However, real-world data often comes from diverse domains, leading to potential mismatches between the domains of synthesized images and those of the evaluation set. Existing methods primarily assume

Cited by 0SourcePDFScholar
2026

Decomposing Prompts, Composing Actions: A Multi-Granularity Prompting Approach for Incremental Action Learning

AAAI 2026technical

Continual learning for action recognition is a critical capability for next-generation Extended Reality (XR) systems. Yet it faces a severe real-world challenge: strict user privacy that prohibits data rehearsal. While recent prompt-based continual learning methods show promise, we argue their core

Cited by 0SourcePDFScholar
2026

Editing Is a Bargaining Game: Balanced Knowledge Editing in Large Language Models

AAAI 2026technical

Large Language Models (LLMs) are prone to generating incorrect or outdated information, thereby necessitating efficient and precise mechanisms for knowledge updates. Existing knowledge editing approaches, however, often encounter conflicts between two competing objectives: maintaining existing knowl

Cited by 0SourcePDFScholar
2026

Hystar: Hypernetwork-driven Style-adaptive Retrieval via Dynamic SVD Modulation

ICLR 2026poster

Query-based image retrieval (QBIR) requires retrieving relevant images given diverse and often stylistically heterogeneous queries, such as sketches, artworks, or low-resolution previews. While large-scale vision--language representation models (VLRMs) like CLIP offer strong zero-shot retrieval perf…

Cited by 0SourceScholar
2026

Learning Attribute–Affordance Hierarchies in Hyperbolic Space for Open-Vocabulary 3D Object Affordance Grounding

ICML 2026poster

This paper pays attention to open-vocabulary 3D object affordance grounding (OVAG), which aims to localize affordance regions on 3D objects by leveraging interaction images or textual instructions. Most existing methods treat interaction images as sources of external affordance knowledge and align t…

Cited by 0SourceScholar
2026

Loc$^{2}$: Interpretable Cross-View Localization via Depth-Lifted Local Feature Matching

ICLR 2026poster

We propose an accurate and interpretable fine-grained cross-view localization method that estimates the 3 Degrees of Freedom (DoF) pose of a ground-level image by matching its local features with a reference aerial image. Unlike prior approaches that rely on global descriptors or bird’s-eye-view (BE…

Cited by 0SourceScholar
2026

Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If Calibrated

AAAI 2026technical

Despite being trained on balanced datasets, existing AI-generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake samples and implicit priors learned during training.

Cited by 0SourcePDFScholar
2025

ChatGarment: Garment Estimation, Generation and Editing via Large Language Models

CVPR 2025poster

We introduce ChatGarment, a novel approach that leverages large vision-language models (VLMs) to automate the estimation, generation, and editing of 3D garment sewing patterns from images or text descriptions. Unlike previous methods that often lack robustness and interactive editing capabilities, C…

Cited by 5SourcePDFScholar
2025

Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative Modeling

NeurIPS 2025poster

In dance performances, choreographers define the visual expression of movement, while cinematographers shape its final presentation through camera work. Consequently, the synthesis of camera movements informed by both music and dance has garnered increasing research interest. While recent advancemen…

Cited by 0SourceScholar
2025

Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization

ICLR 2025poster

Recently, the comprehensive understanding of human motion has been a prominent area of research due to its critical importance in many fields. However, existing methods often prioritize specific downstream tasks and roughly align text and motion features within a CLIP-like framework. This results in…

Cited by 1SourcePDFScholar
2024

Asymmetric Mutual Alignment for Unsupervised Zero-Shot Sketch-Based Image Retrieval

AAAI 2024technical

In recent years, many methods have been proposed to address the zero-shot sketch-based image retrieval (ZS-SBIR) task, which is a practical problem in many applications. However, in real-world scenarios, on the one hand, we can not obtain training data with the same distribution as the test data, an…

Cited by 6SourcePDFScholar
2024

LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures

ACL 2024long

In response to the escalating demand for digital human representations, progress has been made in the generation of realistic human gestures from given speeches. Despite the remarkable achievements of recent research, the generation process frequently includes unintended, meaningless, or non-realist…

Cited by 3SourcePDFScholar
2024

Retrieval Across Any Domains via Large-scale Pre-trained Model

ICML 2024poster

In order to enhance the generalization ability towards unseen domains, universal cross-domain image retrieval methods require a training dataset encompassing diverse domains, which is costly to assemble. Given this constraint, we introduce a novel problem of data-free adaptive cross-domain retrieval…

Cited by 0SourcePDFScholar
2022

Attention-guided Contrastive Hashing for Long-tailed Image Retrieval

IJCAI 2022poster

Image hashing is to represent an image using a binary code for efficient storage and accurate retrieval. Recently, deep hashing methods have shown great improvements on ideally balanced datasets, however, long-tailed data is more common due to rare samples or data collection costs in the real world.…

2022

Noise Is Also Useful: Negative Correlation-Steered Latent Contrastive Learning

CVPR 2022poster

How to effectively handle label noise has been one of the most practical but challenging tasks in Deep Neural Networks (DNNs). Recent popular methods for training DNNs with noisy labels mainly focus on directly filtering out samples with low confidence or repeatedly mining valuable information from…

Cited by 27PDFScholar