Information theoretic clustering for unsupervised domain-adaptation
Subhadeep Dey, Srikanth R. Madikeri, Petr Motlícek
Abstract
The aim of the domain-adaptation task for speaker verification is to exploit unlabelled target domain data by using the labelled source domain data effectively. The i-vector based Probabilistic Linear Discriminant Analysis (PLDA) framework approaches this task by clustering the target domain data and using each cluster as a unique speaker to estimate PLDA model parameters. These parameters are then combined with the PLDA parameters from the source domain. Typically, agglomerative clustering with cosine distance measure is used. In tasks such as speaker diarization that also require unsupervised clustering of speakers, information-theoretic clustering measures have been shown to be effective. In this paper, we employ the Information Bottleneck (IB) clustering technique to find speaker clusters in the target domain data. This is achieved by optimizing the IB criterion that minimizes the information loss during the clustering process. The greedy optimization of the IB criterion involves agglomerative clustering using the Jensen-Shannon divergence as the distance metric. Our experiments in the domain-adaptation task indicate that the proposed system outperforms the baseline by about 14% relative in terms of equal error rate.
BibTeX
@inproceedings{icassp2016_informationtheor,
title = {Information theoretic clustering for unsupervised domain-adaptation},
author = {Subhadeep Dey and Srikanth R. Madikeri and Petr Motlícek},
booktitle = {ICASSP 2016},
year = {2016}
}