← Search

Yazhou Xing

8 accepted papers

2026

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

ICLR 2026poster

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data (e.g., conflating acoustically distinct sounds like different dog barks under co…

Cited by 0SourcecodeScholar
2025

VideoVAE+: Large Motion Video Autoencoding with Cross-modal Video VAE

ICCV 2025poster

Learning a robust video Variational Autoencoder (VAE) is essential for reducing video redundancy and facilitating efficient video generation. Directly applying image VAEs to individual frames in isolation results in temporal inconsistencies and fails to compress temporal redundancy effectively. Exis…

Cited by 0SourcePDFScholar
2024

Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners

CVPR 2024poster

Video and audio content creation serves as the core technique for the movie industry and professional users. Recently existing diffusion-based methods tackle video and audio generation separately which hinders the technique transfer from academia to industry. In this work we aim at filling the gap w…

2019

Leveraging Structural Regularity of Atlanta World for Monocular SLAM

ICRA 2019poster

A wide range of man-made environments can be abstracted as the Atlanta world. It consists of a set of Atlanta frames with a common vertical (gravitational) axis and multiple horizontal axes orthogonal to this vertical axis. This paper focuses on leveraging the regularity of Atlanta world for monocul…

Cited by 48SourceScholar
2018

A Monocular SLAM System Leveraging Structural Regularity in Manhattan World

ICRA 2018poster

The structural features in Manhattan world encode useful geometric information of parallelism, orthogonality and/or coplanarity in the scene. By fully exploiting these structural features, we propose a novel monocular SLAM system which provides accurate estimation of camera poses and 3D map. The for…

Cited by 72SourceScholar