Source cell phone matching from speech recordings by sparse representation and KISS metric
Ling Zou, Qianhua He, Ji-Chen Yang, Yanxiong Li
Abstract
Source recording device matching from two speech recordings is a new and important problem of digital media forensics. It aims to answer the question that whether or not two speech recordings are recorded by the same recording device. In this study we propose a source cell phone matching scheme. The Gaussian supervector (GSV) based on Mel-frequency cepstral coefficients (MFCCs) is extracted from the speech recording and is sparse represented with respect to a dictionary learned by K-SVD algorithm. The reduced-dimensional sparse representation coefficient is utilized to characterize the intrinsic fingerprint of the recording device. Then, KISS metric learning based similarity matching is conducted on a pair of fingerprints extracted from the two speech recordings. Evaluation experiments were conducted on a database of speech recordings recorded by 14 cell phones. The experimental results demonstrated the feasibility of the proposed scheme.
BibTeX
@inproceedings{icassp2016_sourcecellphonem,
title = {Source cell phone matching from speech recordings by sparse representation and KISS metric},
author = {Ling Zou and Qianhua He and Ji-Chen Yang and Yanxiong Li},
booktitle = {ICASSP 2016},
year = {2016}
}