2022
Adversarial Input Ablation for Audio-Visual Learning
ICASSP 2022accepted
We present an adversarial data augmentation strategy for speech spectrograms, within the context of training a model to semantically ground spoken audio captions to the images they describe. Our approach uses a two-pass strategy during training: first, a forward pass through the model is performed i…