← Search

Faisal Ahmed

2 accepted papers

2022

SwinBERT: End-to-End Transformers With Sparse Attention for Video Captioning

CVPR 2022poster

The canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on video frames sampled at a fixed frame rate and are often trained on image/video understanding tasks, without adaption to vi…

Cited by 329PDFcodeScholar
2022

UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling

ECCV 2022poster

"We propose UniTAB that Unifies Text And Box outputs for grounded vision-language (VL) modeling. Grounded VL tasks such as grounded captioning require the model to generate a text description and align predicted words with object regions. To achieve this, models must generate desired text and box ou…