← Search

Yusong Wu

6 accepted papers

2026

Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction

ICLR 2026poster

Most applications of generative AI involve a sequential interaction in which a person inputs a prompt and waits for a response, and where reaction time and adaptivity are not important factors. In contrast, live jamming is a collaborative interaction that requires real-time coordination and adaptati…

Cited by 0SourceScholar
2025

FLAM: Frame-Wise Language-Audio Modeling

ICML 2025poster

Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an…

Cited by 0SourcePDFScholar
2024

Adaptive Accompaniment with ReaLchords

ICML 2024poster

Jamming requires coordination, anticipation, and collaborative creativity between musicians. Current generative models of music produce expressive output but are not able to generate in an online manner, meaning simultaneously with other musicians (human or otherwise). We propose ReaLchords, an onli…

Cited by 3SourcePDFScholar
2024

MusicLDM: Enhancing Novelty in text-to-music Generation Using Beat-Synchronous mixup Strategies

ICASSP 2024accepted

Diffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited availability of music data and sensitive issues related to copyright a…

Cited by 0SourceScholar
2023

Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation

ICASSP 2023accepted

Contrastive learning has shown remarkable success in the field of multimodal representation learning. In this paper, we propose a pipeline of contrastive language-audio pretraining to develop an audio representation by combining audio data with natural language descriptions. To accomplish this targe…

Cited by 0SourceScholar
2022

MIDI-DDSP: Detailed Control of Musical Performance via Hierarchical Modeling

ICLR 2022oral

Musical expression requires control of both what notes that are played, and how they are performed. Conventional audio synthesizers provide detailed expressive controls, but at the cost of realism. Black-box neural audio synthesis and concatenative samplers can produce realistic audio, but have few…