← Search

Luca A Lanzendörfer

11 accepted papers

2026

Text-to-Scene with Large Reasoning Models

AAAI 2026technical

Prompt-driven scene synthesis allows users to generate complete 3D environments from textual descriptions. Current text-to-scene methods often struggle with complex geometries and object transformations, and tend to show weak adherence to complex instructions. We address these limitations by introdu

Cited by 0SourcePDFScholar
2025

Benchmarking Music Generation Models and Metrics via Human Preference Studies

ICASSP 2025accepted

Recent advancements have brought generated music closer to human-created compositions, yet evaluating these models remains challenging. While human preference is the gold standard for assessing quality, translating these subjective judgments into objective metrics, particularly for text-audio alignm…

Cited by 0SourceScholar
2025

Bootstrapping Language-Audio Pre-training for Music Captioning

ICASSP 2025accepted

We introduce BLAP, a model capable of generating high-quality captions for music. BLAP leverages a fine-tuned CLAP audio encoder and a pre-trained Flan-T5 large language model. To achieve effective cross-modal alignment between music and language, BLAP utilizes a Querying Transformer, allowing us to…

Cited by 0SourceScholar
2025

Coarse-to-Fine Text-to-Music Latent Diffusion

ICASSP 2025accepted

We introduce DiscoDiff, a text-to-music generative model that utilizes two latent diffusion models to produce high-fidelity 44.1kHz music hierarchically. Our approach significantly enhances audio quality through a coarse-to-fine generation strategy, leveraging residual vector quantization from the D…

Cited by 0SourceScholar
2025

Contrastive Lyrics Alignment with a Timestamp-Informed Loss

ICASSP 2025accepted

Recent multimodal methods for lyrics alignment have relied on large datasets. Our approach introduces a box loss that directly incorporates timestamp information into the loss function, enabling precise alignment and competitive results even with limited training data. We also address the noise pres…

Cited by 0SourceScholar
2025

EuroSpeech: A Multilingual Speech Corpus

NeurIPS 2025spotlight

Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often contain insufficient data for each language, leading to models tr…

Cited by 0SourceScholar
2025

Generating Vocals from Lyrics and Musical Accompaniment

ICASSP 2025accepted

In this work, we introduce AutoSing, a novel framework designed to generate diverse and high-quality singing voices from provided lyrics and musical accompaniment. AutoSing extends an existing semantic token-based text-to-speech approach by incorporating musical accompaniment as an additional condit…

Cited by 0SourceScholar
2025

High-Fidelity Music Vocoder using Neural Audio Codecs

ICASSP 2025accepted

While neural vocoders have made significant progress in high-fidelity speech synthesis, their application on polyphonic music has remained underexplored. In this work, we propose DisCoder, a neural vocoder that leverages a generative adversarial encoder-decoder architecture informed by a neural audi…

Cited by 0SourceScholar
2025

SAO-Instruct: Free-form Audio Editing using Natural Language Instructions

NeurIPS 2025poster

Generative models have made significant progress in synthesizing high-fidelity audio from short textual descriptions. However, editing existing audio using natural language has remained largely underexplored. Current approaches either require the complete description of the edited audio or are const…

Cited by 0SourceScholar
2024

PUZZLES: A Benchmark for Neural Algorithmic Reasoning

NeurIPS 2024poster

Algorithmic reasoning is a fundamental cognitive ability that plays a pivotal role in problem-solving and decision-making processes. Reinforcement Learning (RL) has demonstrated remarkable proficiency in tasks such as motor control, handling perceptual input, and managing stochastic environments. Th…

2023

DISCO-10M: A Large-Scale Music Dataset

NeurIPS 2023poster

Music datasets play a crucial role in advancing research in machine learning for music. However, existing music datasets suffer from limited size, accessibility, and lack of audio resources. To address these shortcomings, we present DISCO-10M, a novel and extensive music dataset that surpasses the l…

Cited by 18SourcePDFScholar