ICASSP 2025accepted0 citations

Enhancing Unsupervised Acoustic Word Embedding with Visual-Grounded Speech Model and Novel Word-level ABX Evaluation Schemes

Mau Nguyen, Shinobu Hasegawa, Sakriani Sakti

Abstract

Most recent Acoustic Word Embedding (AWE) systems utilize an autoencoder-like approach to compress speech features of arbitrary shapes into fixed-size numerical vectors and then reconstructing it, thereby capturing essential patterns in the data. Unfortunately, AWE models have commonly relied on supervised learning, necessitating extensive textual data, or have employed unsupervised dynamic-based methods that are computationally demanding. This paper introduces an unsupervised approach to AWE, leveraging a self-supervised Visual-Grounded Speech (VGS) model, eliminating the need for dynamic algorithms or textual data. Additionally, we propose a fine-grained ABX evaluation protocol that meticulously assesses the acoustic similarity between spoken segments, providing a more comprehensive and fair evaluation of model performance. Our findings indicate that proposed visual-grounded approach allows the AWE model to function in a truly unsupervised manner without relying on text data and computationally intensive dynamic-based algorithms, while also achieving performance comparable to other approaches.

BibTeX
@inproceedings{icassp2025_enhancingunsuper,
  title = {Enhancing Unsupervised Acoustic Word Embedding with Visual-Grounded Speech Model and Novel Word-level ABX Evaluation Schemes},
  author = {Mau Nguyen and Shinobu Hasegawa and Sakriani Sakti},
  booktitle = {ICASSP 2025},
  year = {2025}
}