← Search

Sebastian Ewert

16 accepted papers

2024

Unsupervised Pitch-Timbre Disentanglement of Musical Instruments Using a Jacobian Disentangled Sequential Autoencoder

ICASSP 2024accepted

Disentangled representation learning seeks to align individual dimensions or separate groups of coordinates of latent factors with attributes of observed data such that perturbing certain latent factors uniquely changes particular attributes. A main challenge in unsupervised disentanglement using au…

Cited by 0SourceScholar
2023

Contrastive Learning-Based Audio to Lyrics Alignment for Multiple Languages

ICASSP 2023accepted

Lyrics alignment gained considerable attention in recent years. State-of-the-art systems either re-use established speech recognition toolkits, or design end-to-end solutions involving a Connectionist Temporal Classification (CTC) loss. However, both approaches suffer from specific weaknesses: toolk…

Cited by 0SourceScholar
2022

A Lightweight Instrument-Agnostic Model for Polyphonic Note Transcription and Multipitch Estimation

ICASSP 2022accepted

Automatic Music Transcription (AMT) has been recognized as a key enabling technology with a wide range of applications. Given the task’s complexity, best results have typically been reported for systems focusing on specific settings, e.g. instrument-specific systems tend to yield improved results ov…

Cited by 0SourceScholar
2022

Towards Robust Unsupervised Disentanglement of Sequential Data — A Case Study Using Music Audio

IJCAI 2022poster

Disentangled sequential autoencoders (DSAEs) represent a class of probabilistic graphical models that describes an observed sequence with dynamic latent variables and a static latent variable. The former encode information at a frame rate identical to the observation, while the latter globally gover…

2020

Seq-U-Net: A One-Dimensional Causal U-Net for Efficient Sequence Modelling

IJCAI 2020poster

Convolutional neural networks (CNNs) with dilated filters such as the Wavenet or the Temporal Convolutional Network (TCN) have shown good results in a variety of sequence modelling tasks. While their receptive field grows exponentially with the number of layers, computing the convolutions over very…

2020

Training Generative Adversarial Networks from Incomplete Observations using Factorised Discriminators

ICLR 2020poster

Generative adversarial networks (GANs) have shown great success in applications such as image generation and inpainting. However, they typically require large datasets, which are often not available, especially in the context of prediction tasks such as image segmentation that require labels. Theref…

Cited by 2SourcecodeScholar
2019

End-to-end Lyrics Alignment for Polyphonic Music Using an Audio-to-character Recognition Model

ICASSP 2019accepted

Time-aligned lyrics can enrich the music listening experience by enabling karaoke, text-based song retrieval and intra-song navigation, and other applications. Compared to text-to-speech alignment, lyrics alignment remains highly challenging, despite many attempts to combine numerous sub-modules inc…

Cited by 0SourceScholar
2018

Adversarial Semi-Supervised Audio Source Separation Applied to Singing Voice Extraction

ICASSP 2018accepted

The state of the art in music source separation employs neural networks trained in a supervised fashion on multi-track databases to estimate the sources from a given mixture. With only few datasets available, often extensive data augmentation is used to combat overfitting. Mixing random tracks, howe…

Cited by 0SourceScholar
2018

Shift-Invariant Kernel Additive Modelling for Audio Source Separation

ICASSP 2018accepted

A major goal in blind source separation to identify and separate sources is to model their inherent characteristics. While most state-of-the-art approaches are supervised methods trained on large datasets, interest in non-data-driven approaches such as Kernel Additive Modelling (KAM) remains high du…

Cited by 0SourceScholar
2017

Improved template based chord recognition using the CRP feature

ICASSP 2017accepted

The task of chord recognition in music signals is often based upon pattern matching in chromagrams. Many variants of chroma exist and quality of chord recognition is related to the feature employed. Chroma Reduced Pitch (CRP) features are interesting in this context as they were designed to improve…

Cited by 0SourceScholar
2017

Interference reduction in music recordings combining Kernel Additive Modelling and Non-Negative Matrix Factorization

ICASSP 2017accepted

In live and studio recordings unexpected sound events often lead to interferences in the signal. For non-stationary interferences, sound source separation techniques can be used to reduce the interference level in the recording. In this context, we present a novel approach combining the strengths of…

Cited by 0SourceScholar
2017

Structured dropout for weak label and multi-instance learning and its application to score-informed source separation

ICASSP 2017accepted

Many success stories involving deep neural networks are instances of supervised learning, where available labels power gradient-based learning methods. Creating such labels, however, can be expensive and thus there is increasing interest in weak labels which only provide coarse information, with unc…

Cited by 0SourceScholar
2016

A score-informed shift-invariant extension of complex matrix factorization for improving the separation of overlapped partials in music recordings

ICASSP 2016accepted

Similar to non-negative matrix factorization (NMF), complex matrix factorization (CMF) can be used to decompose a given music recording into individual sound sources. In contrast to NMF, CMF models both the magnitude and phase of a source, which can improve the separation of overlapped partials. How…

Cited by 0SourceScholar
2015

A dynamic programming variant of non-negative matrix deconvolution for the transcription of struck string instruments

ICASSP 2015accepted

Given a musical audio recording, the goal of music transcription is to determine a score-like representation of the piece underlying the recording. Most current transcription methods employ variants of non-negative matrix factorization (NMF), which often fails to robustly model instruments producing…

Cited by 0SourceScholar
2015

Compensating for asynchronies between musical voices in score-performance alignment

ICASSP 2015accepted

The goal of score-performance synchronisation is to align a given musical score to an audio recording of a performance of the same piece. A major challenge in computing such alignments is to account for musical parameters including the local tempo or playing style. To increase the overall robustness…

Cited by 0SourceScholar