← Search

Mingqian Tang

12 accepted papers

2022

Grow and Merge: A Unified Framework for Continuous Categories Discovery

NeurIPS 2022accept

Although a number of studies are devoted to novel category discovery, most of them assume a static setting where both labeled and unlabeled data are given at once for finding new categories. In this work, we focus on the application scenarios where unlabeled data are continuously fed into the catego…

Cited by 32SourcePDFScholar
2022

Hybrid Relation Guided Set Matching for Few-Shot Action Recognition

CVPR 2022poster

Current few-shot action recognition methods reach impressive performance by learning discriminative features for each video via episodic training and designing various temporal alignment strategies. Nevertheless, they are limited in that (a) learning individual features without considering the entir…

Cited by 121PDFcodeScholar
2022

Learning From Untrimmed Videos: Self-Supervised Video Representation Learning With Hierarchical Consistency

CVPR 2022poster

Natural videos provide rich visual contents for self-supervised learning. Yet most existing approaches for learning spatio-temporal representations rely on manually trimmed videos, leading to limited diversity in visual patterns and limited performance gain. In this work, we aim to learn representat…

Cited by 20PDFScholar
2022

Learning a Condensed Frame for Memory-Efficient Video Class-Incremental Learning

NeurIPS 2022accept

Recent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to the limited memory. To address this problem, we propose FrameMaker, a memory-efficient video class-incremental learning…

Cited by 20SourcePDFScholar
2022

Open-World Semantic Segmentation for LIDAR Point Clouds

ECCV 2022poster

"Classical LIDAR semantic segmentation is not robust for real-world applications, e.g., autonomous driving, since it is closed-set and static. The closed-set network is only able to output labels of trained classes, even for objects never seen before, while a static network cannot update its knowled…

2022

RLIP: Relational Language-Image Pre-training for Human-Object Interaction Detection

NeurIPS 2022accept

The task of Human-Object Interaction (HOI) detection targets fine-grained visual parsing of humans interacting with their environment, enabling a broad range of applications. Prior work has demonstrated the benefits of effective architecture design and integration of relevant cues for more accurate…

2022

Rethinking Supervised Pre-Training for Better Downstream Transferring

ICLR 2022poster

The pretrain-finetune paradigm has shown outstanding performance on many applications of deep learning, where a model is pre-trained on an upstream large dataset (e.g. ImageNet), and is then fine-tuned to different downstream tasks. Though for most cases, the pre-training stage is conducted based on…

Cited by 52SourcePDFScholar
2022

TAda! Temporally-Adaptive Convolutions for Video Understanding

ICLR 2022poster

Spatial convolutions are widely used in numerous deep video models. It fundamentally assumes spatio-temporal invariance, i.e., using shared weights for every location in different frames. This work presents Temporally-Adaptive Convolutions (TAdaConv) for video understanding, which shows that adaptiv…

2021

NGC: A Unified Framework for Learning With Open-World Noisy Data

ICCV 2021poster

The existence of noisy data is prevalent in both the training and testing phases of machine learning systems, which inevitably leads to the degradation of model performance. There have been plenty of works concentrated on learning with in-distribution (IND) noisy labels in the last decade, i.e., som…

Cited by 107PDFScholar
2021

Self-Supervised Motion Learning From Static Images

CVPR 2021poster

Motions are reflected in videos as the movement of pixels, and actions are essentially patterns of inconsistent motions between the foreground and the background. To well distinguish the actions, especially those with complicated spatio-temporal interactions, correctly locating the prominent motion…

Cited by 30PDFcodeScholar
2021

Self-Supervised Video Representation Learning with Constrained Spatiotemporal Jigsaw

IJCAI 2021poster

This paper proposes a novel pretext task for self-supervised video representation learning by exploiting spatiotemporal continuity in videos. It is motivated by the fact that videos are spatiotemporal by nature and a representation learned by detecting spatiotemporal continuity/discontinuity is thus…

Cited by 24SourcePDFScholar
2021

Support-Set Based Cross-Supervision for Video Grounding

ICCV 2021poster

Current approaches for video grounding propose kinds of complex architectures to capture the video-text relations, and have achieved impressive improvements. However, it is hard to learn the complicated multi-modal relations by only architecture designing in fact. In this paper, we introduce a novel…

Cited by 53PDFScholar