ECCV 2022poster127 citations

UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling

Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Faisal Ahmed, Zicheng Liu, Yumao Lu, Lijuan Wang

Abstract

"We propose UniTAB that Unifies Text And Box outputs for grounded vision-language (VL) modeling. Grounded VL tasks such as grounded captioning require the model to generate a text description and align predicted words with object regions. To achieve this, models must generate desired text and box outputs together, and meanwhile indicate the alignments between words and boxes. In contrast to existing solutions that use multiple separate modules for different outputs, UniTAB represents both text and box outputs with a shared token sequence, and introduces a special

BibTeX
@inproceedings{eccv2022_unitabunifyingte,
  title = {UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling},
  author = {Zhengyuan Yang and Zhe Gan and Jianfeng Wang and Xiaowei Hu and Faisal Ahmed and Zicheng Liu and Yumao Lu and Lijuan Wang},
  booktitle = {ECCV 2022},
  year = {2022}
}
UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling · ECCV 2022