← Search

Yihang Zhu

7 accepted papers

2026

Learning Attribute–Affordance Hierarchies in Hyperbolic Space for Open-Vocabulary 3D Object Affordance Grounding

ICML 2026poster

This paper pays attention to open-vocabulary 3D object affordance grounding (OVAG), which aims to localize affordance regions on 3D objects by leveraging interaction images or textual instructions. Most existing methods treat interaction images as sources of external affordance knowledge and align t…

Cited by 0SourceScholar
2026

Multimodal Semantic Bias Mitigation for Diverse Text-To-3D Generation

CVPR 2026

The latest progress in text-to-3D generative models makes it possible to generate high-quality 3D content. Recent text-to-3D large model have achieved remarkable breakthroughs in multi-view consistency. However, their effectiveness is often affected by inherent biases, resulting in sensitivity to de

Cited by 0SourceScholar
2026

Raise One and Infer Three: Toward Reasoning- and Memory-Augmented Diffusion Policy Generalization

IJCAI 2026

Diffusion policy has shown impressive performance in robotic manipulation tasks while struggling with out-of-distribution shifts and limited demonstrations. Recent advances primarily focus on improving geometric or perceptual representations for diffusion policy. However, these approaches rely heavi

Cited by 0Scholar
2025

AffordDP: Generalizable Diffusion Policy with Transferable Affordance

CVPR 2025poster

Diffusion-based policies have shown impressive performance in robotic manipulation tasks while struggling with out-of-domain distributions. Recent efforts attempted to enhance generalization by improving the visual feature encoding for diffusion policy. However, their generalization is typically lim…

Cited by 5SourcePDFScholar
2025

Multi-modal Multi-platform Person Re-Identification: Benchmark and Method

ICCV 2025poster

Conventional person re-identification (ReID) research is often limited to single-modality sensor data from static cameras, which fails to address the complexities of real-world scenarios where multi-modal signals are increasingly prevalent. For instance, consider an urban ReID system integrating sta…

Cited by 0SourcePDFScholar
2025

Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance Grounding

CVPR 2025poster

This paper pays attention to Weakly Supervised Affordance Grounding (WSAG) task that aims to train model to identify affordance regions using human-object interaction images and egocentric images without the need for costly pixel-level annotations. Most existing methods usually consider the affordan…

Cited by 0SourcePDFScholar
2025

VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual Grounding

ICCV 2025poster

As an important direction of embodied intelligence, 3D Visual Grounding has attracted much attention, aiming to identify 3D objects matching the given language description. Most existing methods often follow a two-stage process, i.e., first detecting proposal objects and identifying the right object…

Cited by 0SourcePDFScholar