2024
Siamese Vision Transformers are Scalable Audio-visual Learners
ECCV 2024poster
"Traditional audio-visual methods rely on independent audio and visual backbones, which is costly and not scalable. In this work, we investigate using an audio-visual siamese network () for efficient and scalable audio-visual pretraining. Our framework uses a single shared vision transformer backbon…