← Search

Xiangyuan Yang

2 accepted papers

2026

Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music Generation

AAAI 2026technical

Video-to-music (V2M) generation aims to create music that aligns with visual content. However, two main challenges persist in existing methods: (1) the lack of explicit rhythm modeling hinders audiovisual temporal alignments; (2) effectively integrating various visual features to condition music gen

Cited by 0SourcePDFScholar
2023

A Multi-Stage Triple-Path Method For Speech Separation in Noisy and Reverberant Environments

ICASSP 2023accepted

In noisy and reverberant environments, the performance of deep learning-based speech separation methods drops dramatically because previous methods are not designed and optimized for such situations. To address this issue, we propose a multi-stage end-to-end learning method that decouples the diffic…

Cited by 0SourceScholar