← Search

Yingshan Liang

2 accepted papers

2025

AudioCache: Accelerate Audio Generation With Training-Free Layer Caching

ICASSP 2025accepted

Diffusion models have become the primary choice in audio generation. However, their slow generation speed necessitates acceleration techniques. While current audio generation methods primarily target U-Net-based models, the Diffusion Transformer (DiT) is emerging as the trend in audio generation. As…

Cited by 0SourceScholar
2025

TAGMO: Temporal Control Audio Generation for Multiple Visual Objects Without Training

ICASSP 2025accepted

With the great popularity of Sora, video-based audio generation has become indispensable. While numerous video-to-audio generation models have emerged, they frequently face difficulties including semantic incompatibilities and synchronization problems, especially in situations with multiple objects.…

Cited by 0SourceScholar