← Search

Jiasheng Lu

3 accepted papers

2025

AudioCache: Accelerate Audio Generation With Training-Free Layer Caching

ICASSP 2025accepted

Diffusion models have become the primary choice in audio generation. However, their slow generation speed necessitates acceleration techniques. While current audio generation methods primarily target U-Net-based models, the Diffusion Transformer (DiT) is emerging as the trend in audio generation. As…

Cited by 0SourceScholar
2025

TAGMO: Temporal Control Audio Generation for Multiple Visual Objects Without Training

ICASSP 2025accepted

With the great popularity of Sora, video-based audio generation has become indispensable. While numerous video-to-audio generation models have emerged, they frequently face difficulties including semantic incompatibilities and synchronization problems, especially in situations with multiple objects.…

Cited by 0SourceScholar
2024

BATON: Aligning Text-to-Audio Model Using Human Preference Feedback

IJCAI 2024poster

With the development of AI-Generated Content (AIGC), text-to-audio models are gaining widespread attention. However, it is challenging for these models to generate audio aligned with human preference due to the inherent information density of natural language and limited model understanding ability.…