← Search

Youcai Zhang

3 accepted papers

2024

Tag2Text: Guiding Vision-Language Model via Image Tagging

ICLR 2024poster

This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a…

Cited by 84SourcePDFScholar
2022

On the Efficacy of Small Self-Supervised Contrastive Models without Distillation Signals

AAAI 2022technical

It is a consensus that small models perform quite poorly under the paradigm of self-supervised contrastive learning. Existing methods usually adopt a large off-the-shelf model to transfer knowledge to the small one via distillation. Despite their effectiveness, distillation-based methods may not be…