← Search

Yuhang Yang

16 accepted papers

2026

DanceTogether: Generating Interactive Multi-Person Video without Identity Drifting

ICLR 2026poster

Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first end-to-end diffusion framework that turns a single reference image…

Cited by 0SourceScholar
2026

GIR-Bench: Versatile Benchmark for Generating Images with Reasoning

ICLR 2026poster

Unified multimodal models integrate the reasoning capacity of large language models with both image understanding and generation, showing great promise for advanced multimodal intelligence. However, the community still lacks a rigorous reasoning-centric benchmark to systematically evaluate the align…

Cited by 0SourcecodeScholar
2026

Gloria: Consistent Character Video Generation via Content Anchors

CVPR 2026

Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to preserve identity or leverage non-character-centric information

Cited by 0SourceScholar
2026

TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions

ICLR 2026poster

Hand-object interaction (HOI) is fundamental for humans to express intent. Existing HOI generation research is predominantly confined to fixed grasping patterns, where control is tied to physical priors such as force closure or generic intent instructions, even when expressed through elaborate langu…

Cited by 0SourceScholar
2025

DisPose: Disentangling Pose Guidance for Controllable Human Image Animation

ICLR 2025poster

Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleton pose), recent works have attempted to introduce additional dense conditions (e.g., depth map) to ensure motion alignme…

2025

GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding

CVPR 2025poster

Open-Vocabulary 3D object affordance grounding aims to anticipate "action possibilities" regions on 3D objects with arbitrary instructions, which is crucial for robots to generically perceive real scenarios and respond to operational changes. Existing methods focus on combining images or languages t…

2025

RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis

ACL 2025finding

With the rapid advancement of Large Language Models (LLMs), there is an increasing need for challenging benchmarks to evaluate their capabilities in handling complex tabular data. However, existing benchmarks are either based on outdated data setups or focus solely on simple, flat table structures.…

2025

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference

CVPR 2025poster

While vision-language models like CLIP have shown remarkable success in open-vocabulary tasks, their application is currently confined to image-level tasks, and they still struggle with dense predictions. Recent works often attribute such deficiency in dense predictions to the self-attention layers…

2025

SIGMAN: Scaling 3D Human Gaussian Generation with Millions of Assets

ICCV 2025poster

3D human digitization has long been a highly pursued yet challenging task. Existing methods aim to generate high-quality 3D digital humans from single or multiple views, but remain primarily constrained by current paradigms and the scarcity of 3D human assets. Specifically, recent approaches fall in…

Cited by 0SourcePDFScholar
2024

EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views

NeurIPS 2024poster

Understanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semantics e.g., ''what'' interaction is occurring, capturing ''where'' the interaction…

Cited by 6SourcePDFScholar
2024

LEMON: Learning 3D Human-Object Interaction Relation from 2D Images

CVPR 2024poster

Learning 3D human-object interaction relation is pivotal to embodied AI and interaction modeling. Most existing methods approach the goal by learning to predict isolated interaction elements e.g. human contact object affordance and human-object spatial relation primarily from the perspective of eith…

2023

Grounding 3D Object Affordance from 2D Interactions in Images

ICCV 2023poster

Grounding 3D object affordance seeks to locate objects' "action possibilities" regions in the 3D space, which serves as a link between perception and operation for embodied agents. Existing studies primarily focus on connecting visual affordances with geometry structures, e.g., relying on annotation…

Cited by 34PDFcodeScholar
2023

Hierarchical Softmax for End-To-End Low-Resource Multilingual Speech Recognition

ICASSP 2023accepted

Low-resource speech recognition has been long-suffering from insufficient training data. In this paper, we propose an approach that leverages neighboring languages to improve low-resource scenario performance, founded on the hypothesis that similar linguistic units in neighboring languages exhibit c…

Cited by 0SourceScholar
2023

Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation

ICASSP 2023accepted

Existing speech separation models based on deep learning typically generalize poorly due to domain mismatch. In this paper, we propose SpeakerAugment (SA), a data augmentation method for generalizable speech separation that aims to increase the diversity of speaker identity in training data, to miti…

Cited by 0SourceScholar
2023

Speech-Text Based Multi-Modal Training with Bidirectional Attention for Improved Speech Recognition

ICASSP 2023accepted

To let the state-of-the-art end-to-end ASR model enjoy data efficiency, as well as much more unpaired text data by multi-modal training, one needs to address two problems: 1) the synchronicity of feature sampling rates between speech and language (aka text data); 2) the homogeneity of the learned re…

Cited by 0SourceScholar