2024
Speech Guided Masked Image Modeling for Visually Grounded Speech
ICASSP 2024accepted
The objective of this study is to investigate the learning process of Visually Grounded Speech (VGS) models through joint learning that combines contrastive learning and masked image modeling. Typically, VGS models ahn to establish audio-visual alignment between images and then spoken captions withi…