← Search

Yin-Jyun Luo

5 accepted papers

2024

Posterior Variance-Parameterised Gaussian Dropout: Improving Disentangled Sequential Autoencoders for Zero-Shot Voice Conversion

ICASSP 2024accepted

The class of disentangled sequential auto-encoders factorises speech into time-invariant (global) and time-variant (local) representations for speaker identity and linguistic content, respectively. Many of the existing models employ this assumption to tackle zero-shot voice conversion (VC), which co…

Cited by 0SourceScholar
2024

Unsupervised Pitch-Timbre Disentanglement of Musical Instruments Using a Jacobian Disentangled Sequential Autoencoder

ICASSP 2024accepted

Disentangled representation learning seeks to align individual dimensions or separate groups of coordinates of latent factors with attributes of observed data such that perturbing certain latent factors uniquely changes particular attributes. A main challenge in unsupervised disentanglement using au…

Cited by 0SourceScholar
2022

Towards Robust Unsupervised Disentanglement of Sequential Data — A Case Study Using Music Audio

IJCAI 2022poster

Disentangled sequential autoencoders (DSAEs) represent a class of probabilistic graphical models that describes an observed sequence with dynamic latent variables and a static latent variable. The former encode information at a frame rate identical to the observation, while the latter globally gover…

2020

Singing Voice Conversion with Disentangled Representations of Singer and Vocal Technique Using Variational Autoencoders

ICASSP 2020accepted

We propose a flexible framework that deals with both singer conversion and singers vocal technique conversion. The proposed model is trained on non-parallel corpora, accommodates many-to-many conversion, and leverages recent advances of variational autoencoders. It employs separate encoders to learn…

Cited by 0SourceScholar