2025
Enhancing Unsupervised Acoustic Word Embedding with Visual-Grounded Speech Model and Novel Word-level ABX Evaluation Schemes
ICASSP 2025accepted
Most recent Acoustic Word Embedding (AWE) systems utilize an autoencoder-like approach to compress speech features of arbitrary shapes into fixed-size numerical vectors and then reconstructing it, thereby capturing essential patterns in the data. Unfortunately, AWE models have commonly relied on sup…