← Search

Wei-Hsiang Liao

17 accepted papers

2026

Concept-TRAK: Understanding how diffusion models learn concepts through concept attribution

ICLR 2026poster

While diffusion models excel at image generation, their growing adoption raises critical concerns about copyright issues and model transparency. Existing attribution methods identify training examples influencing an entire image, but fall short in isolating contributions to specific elements, such a…

Cited by 0SourcecodeScholar
2026

LLM2Fx-Tools: Tool Calling for Music Post-Production

ICLR 2026poster

This paper introduces LLM2Fx-Tools, a multimodal tool-calling framework that generates executable sequences of audio effects (Fx-chain) for music post-production. LLM2Fx-Tools uses a large language model (LLM) to understand audio inputs, select audio effects types, determine their order, and estimat…

Cited by 0SourcecodeScholar
2026

SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music Editing

AAAI 2026technical

Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing methods rely on pretrained diffusion models by involving forward-backward diffusion processes. However, these methods ofte

Cited by 0SourcePDFScholar
2025

HERO: Human-Feedback Efficient Reinforcement Learning for Online Diffusion Model Finetuning

ICLR 2025poster

Controllable generation through Stable Diffusion (SD) fine-tuning aims to improve fidelity, safety, and alignment with human guidance. Existing reinforcement learning from human feedback methods usually rely on predefined heuristic reward functions or pretrained reward models built on large-scale da…

2025

Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer

ICASSP 2025accepted

Music timbre transfer is a challenging task that involves modifying the timbral characteristics of an audio signal while preserving its melodic structure. In this paper, we propose a novel method based on dual diffusion bridges, trained using the CocoChorales Dataset, which consists of unpaired mono…

Cited by 0SourceScholar
2025

Variable Bitrate Residual Vector Quantization for Audio Coding

ICASSP 2025accepted

Recent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be suboptimal in terms of rate-distortion tradeoff, particularly in scenarios with simple…

Cited by 12SourceScholar
2024

Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion

ICLR 2024poster

Consistency Models (CM) (Song et al., 2023) accelerate score-based diffusion model sampling at the cost of sample quality but lack a natural way to trade-off quality for speed. To address this limitation, we propose Consistency Trajectory Model (CTM), a generalization encompassing CM and score-based…

2024

Manifold Preserving Guided Diffusion

ICLR 2024poster

Despite the recent advancements, conditional image generation still faces challenges of cost, generalizability, and the need for task-specific training. In this paper, we propose Manifold Preserving Guided Diffusion (MPGD), a training-free conditional generation framework that leverages pretrained d…

Cited by 50SourcePDFScholar
2024

MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models

IJCAI 2024poster

Recent advances in text-to-music generation models have opened new avenues in musical creativity. However, the task of editing these generated music remains a significant challenge. This paper introduces a novel approach to edit music generated by such models, enabling the modification of specific a…

2024

On the Language Encoder of Contrastive Cross-modal Models

ACL 2024findings

Contrastive cross-modal models such as CLIP and CLAP aid various vision-language (VL) and audio-language (AL) tasks. However, there has been limited investigation of and improvement in their language encoder – the central component of encoding natural language descriptions of image/audio into vector…

Cited by 0SourcePDFScholar
2024

PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion Teacher

NeurIPS 2024poster

The diffusion model performs remarkable in generating high-dimensional content but is computationally intensive, especially during training. We propose Progressive Growing of Diffusion Autoencoder (PaGoDA), a novel pipeline that reduces the training costs through three stages: training diffusion on…

2024

Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription

ICASSP 2024accepted

In recent years, research on music transcription has focused mainly on architecture design and instrument-specific data acquisition. With the lack of availability of diverse datasets, progress is often limited to solo-instrument tasks such as piano transcription. Several works have explored multi-in…

Cited by 0SourceScholar
2024

VRDMG: Vocal Restoration via Diffusion Posterior Sampling with Multiple Guidance

ICASSP 2024accepted

Restoring degraded music signals is essential to enhance audio quality for downstream music manipulation. Recent diffusion-based music restoration methods have demonstrated impressive performance, and among them, diffusion posterior sampling (DPS) stands out given its intrinsic properties, making it…

Cited by 0SourceScholar
2023

GNN-Based Point Cloud Maps Feature Extraction and Residual Feature Fusion for 3D Object Detection

ICRA 2023poster

LiDAR detection of long-range vehicles is challenging because very few and sparse points are measured in long distances and vehicles with similar shapes of targets could lead to false positives easily. To tackle these challenges, taking the environment information (HD maps) into account could be ben…

Cited by 2SourceScholar
2023

Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects

ICASSP 2023accepted

We propose an end-to-end music mixing style transfer system that converts the mixing style of an input multitrack to that of a reference song. This is achieved with an encoder pre-trained with a contrastive objective to extract only audio effects related information from a reference music recording.…

Cited by 0SourceScholar
2022

Automatic DJ Transitions with Differentiable Audio Effects and Generative Adversarial Networks

ICASSP 2022accepted

A central task of a Disc Jockey (DJ) is to create a mixset of music with seamless transitions between adjacent tracks. In this paper, we explore a data-driven approach that uses a generative adversarial network to create the song transition by learning from real-world DJ mixes. The generator uses tw…

Cited by 0SourceScholar
2022

Reconstruction and Synthesis of Lidar Point Clouds of Spray

RA-L 2022

Lidars are commonly used on autonomous vehicles, but their performance can be significantly affected by adverse weather. A number of studies have been devoted to analyzing and improving lidars’ performance in rain, fog, and snow. Yet, relatively little attention has been paid to road spray which occ

Cited by 14SourceScholar