2024
Curriculum Masking in Vision-Language Pretraining to Maximize Cross Modal Interaction
NAACL 2024long
Many leading methods in Vision and language (V+L) pretraining utilize masked language modeling (MLM) as a standard pretraining component, with the expectation that reconstruction of masked text tokens would necessitate reference to corresponding image context via cross/self attention and thus promot…