← Search

Xihua Wang

6 accepted papers

2025

Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies

NeurIPS 2025poster

Existing imitation learning methods decouple perception and action, which overlooks the causal reciprocity between sensory representations and action execution that humans naturally leverage for adaptive behaviors. To bridge this gap, we introduce Action-Guided Diffusion Policy (DP-AG), a unified re…

Cited by 0SourcecodeScholar
2025

Enhancing Audiovisual Speech Recognition Through Bifocal Preference Optimization

AAAI 2025technical

Audiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-world scenarios across various domains due to noisy acoustic environments, spontaneous speech, and the uncertain use of vis…

2025

Two-in-One: Unified Multi-Person Interactive Motion Generation by Latent Diffusion Transformer

ICASSP 2025accepted

Multi-person interactive motion generation, a critical yet under-explored domain in computer character animation, poses significant challenges such as intricate modeling of inter-human interactions beyond individual motions and generating two motions with huge differences from one text condition. Cu…

Cited by 0SourceScholar
2025

VAFlow: Video-to-Audio Generation with Cross-Modality Flow Matching

ICCV 2025poster

Video-to-audio (V2A) generation aims to synthesize temporally aligned, realistic sounds for silent videos, a critical capability for immersive multimedia applications. Current V2A methods, predominantly based on diffusion or flow models, rely on suboptimal noise-to-audio paradigms that entangle cros…