← Search

Xindi Shang

5 accepted papers

2023

Meta Compositional Referring Expression Segmentation

CVPR 2023poster

Referring expression segmentation aims to segment an object described by a language expression from an image. Despite the recent progress on this task, existing models tackling this task may not be able to fully capture semantics and visual representations of individual concepts, which limits their…

Cited by 34SourcePDFScholar
2023

Token Boosting for Robust Self-Supervised Visual Transformer Pre-Training

CVPR 2023poster

Learning with large-scale unlabeled data has become a powerful tool for pre-training Visual Transformers (VTs). However, prior works tend to overlook that, in real-world scenarios, the input data may be corrupted and unreliable. Pre-training VTs on such corrupted data can be challenging, especially…

Cited by 6SourcePDFScholar
2021

NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions

CVPR 2021poster

We introduce NExT-QA, a rigorously designed video question answering (VideoQA) benchmark to advance video understanding from describing to explaining the temporal actions. Based on the dataset, we set up multi-choice and open-ended QA tasks targeting at causal action reasoning, temporal action reaso…

Cited by 463PDFcodeScholar
2016

Online Collaborative Learning for Open-Vocabulary Visual Classifiers

CVPR 2016poster

We focus on learning open-vocabulary visual classifiers, which scale up to a large portion of natural language vocabulary (e.g., over tens of thousands of classes). In particular, the training data are large-scale weakly labeled Web images since it is difficult to acquire sufficient well-labeled dat…

Cited by 54PDFScholar