2024
FuseTeacher: Modality-fused Encoders are Strong Vision Supervisors
ECCV 2024poster
"Learning visual representation with image-text datasets attracts a lot of attention in recent years. Existing approaches primarily rely on cross-modality supervision, and incorporate intra-modality supervision if necessary. They overlook the potential benefits of modality-fused supervision. Since m…