Speaker segmentation using i-vector in meetings domain
Leonardo Valeriano Neri, Hector N. B. Pinheiro, Tsang Ing Ren, George D. C. Cavalcanti, André Gustavo Adami
Abstract
In this paper, we propose a speaker segmentation method for meeting audio based on i-vector. The motivation is to utilize the Total Variability (TV) framework as a feature extractor and to exploit the potential of modeling the speaker and channel variabilities for speaker segmentation in meetings. A distance-based segmentation method is designed with the cosine distance. A sliding window with variable length searches for speaker turns, through the distance between the i-vectors extracted from two segments with the same size. The experiments are conducted on the AMI Meeting Corpus, covering several conversation scenarios. For the training data of the UBM and TV matrix, 5 conversations from AMI Meeting Corpus are sampled. Other 10 conversations from AMI Meeting Corpus to compose the test data. The experiments show an improvement in the MDR and FAR curves compared with the FixSlid approach with different distance metrics, and for most of the operating points when compared with the classical BIC based WinGrow. The proposed method has on average a better computational performance, improving in 61.5% compared with the XBIC based FixSlid, and improving in 86.7% compared with the BIC based WinGrow.
BibTeX
@inproceedings{icassp2017_speakersegmentat,
title = {Speaker segmentation using i-vector in meetings domain},
author = {Leonardo Valeriano Neri and Hector N. B. Pinheiro and Tsang Ing Ren and George D. C. Cavalcanti and André Gustavo Adami},
booktitle = {ICASSP 2017},
year = {2017}
}