ICASSP 2019accepted0 citations

A Factorial Deep Markov Model for Unsupervised Disentangled Representation Learning from Speech

Sameer Khurana, Shafiq Rayhan Joty, Ahmed Ali, James R. Glass

Abstract

We present the Factorial Deep Markov Model (FDMM) for representation learning of speech. The FDMM learns disentangled, interpretable and lower dimensional latent representations from speech without supervision. We use a static and dynamic latent variable to exploit the fact that information in a speech signal evolves at different time scales. Latent representations learned by the FDMM outperform a baseline i-vector system on speaker verification and dialect identification while also reducing the error rate of a phone recognition system in a domain mismatch scenario.

BibTeX
@inproceedings{icassp2019_afactorialdeepma,
  title = {A Factorial Deep Markov Model for Unsupervised Disentangled Representation Learning from Speech},
  author = {Sameer Khurana and Shafiq Rayhan Joty and Ahmed Ali and James R. Glass},
  booktitle = {ICASSP 2019},
  year = {2019}
}
A Factorial Deep Markov Model for Unsupervised Disentangled Representation Learning from Speech · ICASSP 2019