← Search

Dorien Herremans

9 accepted papers

2026

Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social Interactions

ICLR 2026poster

Large language models (LLMs) are increasingly deployed in multi-agent systems (MAS) as components of collaborative intelligence, where peer interactions dynamically shape individual decision-making. While prior work has largely focused on conformity bias, we broaden the scope to examine how LLMs bui…

Cited by 0SourceScholar
2026

SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering

ICML 2026poster

Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment or expertise. These problems are typically corrected using …

Cited by 0SourcecodeScholar
2025

Coarse-to-Fine Text-to-Music Latent Diffusion

ICASSP 2025accepted

We introduce DiscoDiff, a text-to-music generative model that utilizes two latent diffusion models to produce high-fidelity 44.1kHz music hierarchically. Our approach significantly enhances audio quality through a coarse-to-fine generation strategy, leveraging residual vector quantization from the D…

Cited by 0SourceScholar
2025

Text2midi: Generating Symbolic Music from Captions

AAAI 2025technical

This paper introduces text2midi, an end-to-end model to generate MIDI files from textual descriptions. Leveraging the growing popularity of multimodal generative approaches, text2midi capitalizes on the extensive availability of textual data and the success of large language models (LLMs). Our end-t…

2024

Mustango: Toward Controllable Text-to-Music Generation

NAACL 2024long

The quality of the text-to-music models has reached new heights due to recent advancements in diffusion models. The controllability of various musical aspects, however, has barely been explored. In this paper, we propose Mustango: a music-domain-knowledge-inspired text-to-music system based on diffu…

2023

A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music Modeling

AAAI 2023technical

Following the success of the transformer architecture in the natural language domain, transformer-like architectures have been widely applied to the domain of symbolic music recently. Symbolic music and text, however, are two different modalities. Symbolic music contains multiple attributes, both ab…

2023

Diffroll: Diffusion-Based Generative Music Transcription with Unsupervised Pretraining Capability

ICASSP 2023accepted

In this paper we propose a novel generative approach, DiffRoll, to tackle automatic music transcription (AMT). Instead of treating AMT as a discriminative task in which the model is trained to convert spectrograms into piano rolls, we think of it as a conditional generative task where we train our m…

Cited by 0SourceScholar
2020

Singing Voice Conversion with Disentangled Representations of Singer and Vocal Technique Using Variational Autoencoders

ICASSP 2020accepted

We propose a flexible framework that deals with both singer conversion and singers vocal technique conversion. The proposed model is trained on non-parallel corpora, accommodates many-to-many conversion, and leverages recent advances of variational autoencoders. It employs separate encoders to learn…

Cited by 0SourceScholar