← Search

Xiaoyun Yang

13 accepted papers

2024

Dynamic Semantic-Based Spatial Graph Convolution Network for Skeleton-Based Human Action Recognition

AAAI 2024technical

Graph convolutional networks (GCNs) have attracted great attention and achieved remarkable performance in skeleton-based action recognition. However, most of the previous works are designed to refine skeleton topology without considering the types of different joints and edges, making them infeasibl…

2022

DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image Classification

CVPR 2022oral

Multiple instance learning (MIL) has been increasingly used in the classification of histopathology whole slide images (WSIs). However, MIL approaches for this specific classification problem still face unique challenges, particularly those related to small sample cohorts. In these, there are limite…

Cited by 407PDFcodeScholar
2021

Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box Estimation

CVPR 2021poster

Visual object tracking aims to precisely estimate the bounding box for the given target, which is a challenging problem due to factors such as deformation and occlusion. Many recent trackers adopt the multiple-stage tracking strategy to improve the quality of bounding box estimation. These methods f…

Cited by 268PDFcodeScholar
2021

Spatial Uncertainty-Aware Semi-Supervised Crowd Counting

ICCV 2021poster

Semi-supervised approaches for crowd counting attract attention, as the fully supervised paradigm is expensive and laborious due to its request for a large number of images of dense crowd scenarios and their annotations. This paper proposes a spatial uncertainty-aware semi-supervised approach via re…

Cited by 123PDFcodeScholar
2021

Video Annotation for Visual Tracking via Selection and Refinement

ICCV 2021poster

Deep learning based visual trackers entail offline pre-training on large volumes of video datasets with accurate bounding box annotations that are labor-expensive to achieve. We present a new framework to facilitate bounding box annotations for video sequences, which investigates a selection-and-ref…

Cited by 11PDFcodeScholar
2021

Watching You: Global-Guided Reciprocal Learning for Video-Based Person Re-Identification

CVPR 2021poster

Video-based person re-identification (Re-ID) aims to automatically retrieve video sequences of the same person under non-overlapping cameras. To achieve this goal, it is the key to fully utilize abundant spatial and temporal cues in videos. Existing methods usually focus on the most conspicuous imag…

Cited by 118PDFcodeScholar
2020

Cooling-Shrinking Attack: Blinding the Tracker With Imperceptible Noises

CVPR 2020poster

Adversarial attack of CNN aims at deceiving models to misbehave by adding imperceptible perturbations to images. This feature facilitates to understand neural networks deeply and to improve the robustness of deep learning models. Although several works have focused on attacking image classifiers and…

Cited by 105PDFcodeScholar
2020

High-Performance Long-Term Tracking With Meta-Updater

CVPR 2020oral

Long-term visual tracking has drawn increasing attention because it is much closer to practical applications than short-term tracking. Most top-ranked long-term trackers adopt the offline-trained Siamese architectures, thus,they cannot benefit from great progress of short-term trackers with online u…

Cited by 319PDFcodeScholar
2020

Regression of Instance Boundary by Aggregated CNN and GCN

ECCV 2020poster

This paper proposes a straightforward, intuitive deep learning approach for (biomedical) image segmentation tasks. Different from the existing dense pixel classification methods, we develop a novel multilevel aggregation network to directly regress the coordinates of the boundary of instances in an…

Cited by 33SourcePDFScholar
2019

'Skimming-Perusal' Tracking: A Framework for Real-Time and Robust Long-Term Tracking

ICCV 2019poster

Compared with traditional short-term tracking, long-term tracking poses more challenges and is much closer to realistic applications. However, few works have been done and their performance have also been limited. In this work, we present a novel robust and real-time long-term tracking framework bas…

Cited by 229PDFcodeScholar
2019

Cascaded Context Pyramid for Full-Resolution 3D Semantic Scene Completion

ICCV 2019oral

Semantic Scene Completion (SSC) aims to simultaneously predict the volumetric occupancy and semantic category of a 3D scene. It helps intelligent devices to understand and interact with the surrounding scenes. Due to the high-memory requirement, current methods only produce low-resolution completion…

Cited by 77PDFScholar
2019

GradNet: Gradient-Guided Network for Visual Object Tracking

ICCV 2019oral

The fully-convolutional siamese network based on template matching has shown great potentials in visual tracking. During testing, the template is fixed with the initial target feature and the performance totally relies on the general matching ability of the siamese network. However, this manner cann…

Cited by 400PDFcodeScholar