← Search

Sanna Wager

5 accepted papers

2022

Upmixing Via Style Transfer: A Variational Autoencoder for Disentangling Spatial Images And Musical Content

ICASSP 2022accepted

In the stereo-to-multichannel upmixing problem for music, one of the main tasks is to set the directionality of the instrument sources in the multichannel rendering results. In this paper, we propose a modified variational autoencoder model that learns a latent space to describe the spatial images i…

Cited by 0SourceScholar
2020

Deep Autotuner: A Pitch Correcting Network for Singing Performances

ICASSP 2020accepted

We introduce a data-driven approach to automatic pitch correction of solo singing performances. The proposed approach predicts note-wise pitch shifts from the relationship between the respective spectrograms of the singing and accompaniment. This approach differs from commercial systems, where vocal…

Cited by 0SourceScholar
2020

Fully Learnable Front-End for Multi-Channel Acoustic Modeling Using Semi-Supervised Learning

ICASSP 2020accepted

In this work, we investigated the teacher-student training paradigm to train a fully learnable multi-channel acoustic model for far-field automatic speech recognition (ASR). Using a large offline teacher model trained on beamformed audio, we trained a simpler multi-channel student acoustic model use…

Cited by 0SourceScholar
2019

Intonation: A Dataset of Quality Vocal Performances Refined by Spectral Clustering on Pitch Congruence

ICASSP 2019accepted

We introduce the "Intonation" dataset of amateur vocal performances with a tendency for good intonation, collected from Smule, Inc. The dataset can be used for music information retrieval tasks such as autotuning, query by humming, and singing style analysis. It is available upon request on the Stan…

Cited by 8SourceScholar
2017

Towards expressive instrument synthesis through smooth frame-by-frame reconstruction: From string to woodwind

ICASSP 2017accepted

We consider the task of mapping the performance of a musical excerpt on one instrument to another. Our focus is on excitation-continuous instruments, where pitch, amplitude, spectrum, and time envelope are controlled continuously by the player. The synthesized instrument should follow the target ins…

Cited by 0SourceScholar