← Search

Florian Grötschla

9 accepted papers

2026

SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

ICML 2026poster

We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio seque…

Cited by 0SourceScholar
2025

Benchmarking Music Generation Models and Metrics via Human Preference Studies

ICASSP 2025accepted

Recent advancements have brought generated music closer to human-created compositions, yet evaluating these models remains challenging. While human preference is the gold standard for assessing quality, translating these subjective judgments into objective metrics, particularly for text-audio alignm…

Cited by 0SourceScholar
2025

Contrastive Lyrics Alignment with a Timestamp-Informed Loss

ICASSP 2025accepted

Recent multimodal methods for lyrics alignment have relied on large datasets. Our approach introduces a box loss that directly incorporates timestamp information into the loss function, enabling precise alignment and competitive results even with limited training data. We also address the noise pres…

Cited by 0SourceScholar
2025

EuroSpeech: A Multilingual Speech Corpus

NeurIPS 2025spotlight

Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often contain insufficient data for each language, leading to models tr…

Cited by 0SourceScholar
2025

Generating Vocals from Lyrics and Musical Accompaniment

ICASSP 2025accepted

In this work, we introduce AutoSing, a novel framework designed to generate diverse and high-quality singing voices from provided lyrics and musical accompaniment. AutoSing extends an existing semantic token-based text-to-speech approach by incorporating musical accompaniment as an additional condit…

Cited by 0SourceScholar
2025

High-Fidelity Music Vocoder using Neural Audio Codecs

ICASSP 2025accepted

While neural vocoders have made significant progress in high-fidelity speech synthesis, their application on polyphonic music has remained underexplored. In this work, we propose DisCoder, a neural vocoder that leverages a generative adversarial encoder-decoder architecture informed by a neural audi…

Cited by 0SourceScholar
2025

SAO-Instruct: Free-form Audio Editing using Natural Language Instructions

NeurIPS 2025poster

Generative models have made significant progress in synthesizing high-fidelity audio from short textual descriptions. However, editing existing audio using natural language has remained largely underexplored. Current approaches either require the complete description of the edited audio or are const…

Cited by 0SourceScholar
2024

CoRe-GD: A Hierarchical Framework for Scalable Graph Visualization with GNNs

ICLR 2024poster

Graph Visualization, also known as Graph Drawing, aims to find geometric embeddings of graphs that optimize certain criteria. Stress is a widely used metric; stress is minimized when every pair of nodes is positioned at their shortest path distance. However, stress optimization presents computationa…

2023

DISCO-10M: A Large-Scale Music Dataset

NeurIPS 2023poster

Music datasets play a crucial role in advancing research in machine learning for music. However, existing music datasets suffer from limited size, accessibility, and lack of audio resources. To address these shortcomings, we present DISCO-10M, a novel and extensive music dataset that surpasses the l…

Cited by 18SourcePDFScholar