← Search

Hao-Wen Dong

9 accepted papers

2025

FUTGA-MIR: Enhancing Fine-grained and Temporally-aware Music Understanding with Music Information Retrieval

ICASSP 2025accepted

Recent music large language models (music LLMs) have shown great potential in music understanding through large-scale multimodal pre-training. While some existing music LLMs have been augmented with temporally-aware music captions, music information retrieval (MIR) features conventionally do not exi…

Cited by 0SourceScholar
2025

REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing

NeurIPS 2025poster

Short videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coherent narrative, existing abstractive methods cannot `quote' from the input videos, i.e., inserting short video clips in the…

Cited by 0SourceScholar
2025

Synthesizing Composite Hierarchical Structure from Symbolic Music Corpora

IJCAI 2025

Western music is an innately hierarchical system of interacting levels of structure, from fine-grained melody to high-level form. In order to analyze music compositions holistically and at multiple granularities, we propose a unified, hierarchical meta-representation of musical structure called the

2025

TeaserGen: Generating Teasers for Long Documentaries

ICLR 2025poster

Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling capability for the input videos, while necessitating maintaining audiovisual al…

Cited by 0SourcePDFScholar
2025

ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning

ICASSP 2025accepted

Modeling the natural contour of fundamental frequency (F0) plays a critical role in music audio synthesis. However, transcribing and managing multiple F0 contours in polyphonic music is challenging, and explicit F0 contour modeling has not yet been explored for polyphonic instrumental synthesis. In…

Cited by 0SourceScholar
2023

CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos

ICLR 2023poster

Recent years have seen progress beyond domain-specific sound separation for speech or music towards universal sound separation for arbitrary sounds. Prior work on universal sound separation has investigated separating a target sound out of an audio mixture given a text query. Such text-queried sound…

2022

Deep Performer: Score-to-Audio Music Performance Synthesis

ICASSP 2022accepted

Music performance synthesis aims to synthesize a musical score into a natural performance. In this paper, we borrow recent advances in text-to-speech synthesis and present the Deep Performer—a novel system for score-to-audio music performance synthesis. Unlike speech, music often contains polyphony…

Cited by 0SourceScholar