← Search

Bing Shuai

18 accepted papers

2024

Robustness Preserving Fine-tuning using Neuron Importance

ECCV 2024poster

"Robust fine-tuning aims to adapt a vision-language model to downstream tasks while preserving its zero-shot capabilities on unseen data. Recent studies have introduced fine-tuning strategies to improve in-distribution (ID) performance on the downstream tasks while minimizing deterioration in out-of…

Cited by 0SourcePDFScholar
2024

Self-Supervised Multi-Object Tracking with Path Consistency

CVPR 2024highlight

In this paper we propose a novel concept of path consistency to learn robust object matching without using manual object identity supervision. Our key idea is that to track a object through frames we can obtain multiple different association results from a model by varying the frames it can observe…

2023

MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation

ICCV 2023poster

Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently processing long-form videos (>60min). In this paper, we introduce Mu…

Cited by 4PDFcodeScholar
2023

SkeleTR: Towards Skeleton-based Action Recognition in the Wild

ICCV 2023poster

We present SkeleTR, a new framework for skeleton-based action recognition. In contrast to prior work, which focuses mainly on controlled environments, we target in-the-wild scenarios that typically involve a variable number of people and various forms of interaction between people. SkeleTR works wit…

Cited by 35PDFScholar
2022

An In-depth Study of Stochastic Backpropagation

NeurIPS 2022accept

In this paper, we provide an in-depth study of Stochastic Backpropagation (SBP) when training deep neural networks for standard image classification and object detection tasks. During backward propagation, SBP calculates gradients by using only a subset of feature maps to save GPU memory and computa…

2022

Large Scale Real-World Multi-person Tracking

ECCV 2022poster

"This paper presents a new large scale multi-person tracking dataset. Our dataset is over an order of magnitude larger than currently available high quality multi-object tracking datasets such as MOT17, HiEve, and MOT20 datasets. The lack of large scale training and test data for this task has limit…

2022

TubeR: Tubelet Transformer for Video Action Detection

CVPR 2022oral

We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an off-line actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in video by simulta…

Cited by 95PDFScholar
2019

Semantic Correlation Promoted Shape-Variant Context for Segmentation

CVPR 2019oral

Context is essential for semantic segmentation. Due to the diverse shapes of objects and their complex layout in various scene images, the spatial scales and shapes of contexts for different objects have very large variation. It is thus ineffective or inefficient to aggregate various context informa…

Cited by 215PDFcodeScholar
2018

Context Contrasted Feature and Gated Multi-Scale Aggregation for Scene Segmentation

CVPR 2018poster

Scene segmentation is a challenging task as it need label every pixel in the image. It is crucial to exploit discriminative context and aggregate multi-scale features to achieve better segmentation. In this paper, we first propose a novel context contrasted local feature that not only leverages the…

Cited by 450SourcePDFScholar
2017

Episodic CAMN: Contextual Attention-Based Memory Networks With Iterative Feedback for Scene Labeling

CVPR 2017poster

Scene labeling can be seen as a sequence-sequence prediction task (pixels-labels), and it is quite important to leverage relevant context to enhance the performance of pixel classification. In this paper, we introduce an episodic attention-based memory network to achieve the goal. We present a unifi…

Cited by 17PDFScholar
2015

Integrating Parametric and Non-Parametric Models For Scene Labeling

CVPR 2015poster

We adopt Convolutional Neural Networks (CNN) as our parametric model to learn discriminative features and classifiers for local patch classification. As visually similar pixels are indistinguishable from local context, we alleviate such ambiguity by putting a global scene constraint. We estimate the…

Cited by 56SourcePDFScholar