← Search

Ling Shao

164 accepted papers

2026

MuSASplat: Efficient Sparse-View 3D Gaussian Splats via Lightweight Multi-Scale Adaptation

AAAI 2026technical

Sparse-view 3D Gaussian splatting seeks to render high-quality novel views of 3D scenes from a limited set of input images. While recent pose-free feed-forward methods leveraging pre-trained 3D priors have achieved impressive results, most of them rely on full fine-tuning of large Vision Transformer

Cited by 0SourcePDFScholar
2025

MambaIC: State Space Models for High-Performance Learned Image Compression

CVPR 2025poster

A high-performance image compression algorithm is crucial for real-time information transmission across numerous fields. Despite rapid progress in image compression, computational inefficiency and poor redundancy modeling still pose significant bottlenecks, limiting practical applications. Inspired…

2025

PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations

ICCV 2025poster

COLMAP-free 3D Gaussian Splatting (3D-GS) has recently attracted increasing attention due to its remarkable performance in reconstructing high-quality 3D scenes from unposed images or videos. However, it often struggles to handle scenes with complex camera trajectories as featured by drastic rotatio…

Cited by 0SourcePDFScholar
2025

Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data

ICCV 2025poster

AI for tumor segmentation is limited by the lack of large, voxel-wise annotated datasets, which are hard to create and require medical experts. In our proprietary JHH dataset of 3,000 annotated pancreatic tumor scans, we found that AI performance stopped improving after 1,500 scans. With synthetic d…

2025

Spatial Preference Rewarding for MLLMs Spatial Understanding

ICCV 2025poster

Multimodal large language models(MLLMs) have demonstrated promising spatial understanding capabilities, such as referencing and grounding object descriptions. Despite their successes, MLLMs still fall short in fine-grained spatial perception abilities, such as generating detailed region descriptions…

2025

Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting

ICCV 2025poster

Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of large-scale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and laborious annotation. In this work, we present a scalable pipeline…

2024

CAT-SAM: Conditional Tuning for Few-Shot Adaptation of Segment Anything Model

ECCV 2024oral

"The Segment Anything Model (SAM) has demonstrated remarkable zero-shot capability and flexible geometric prompting in general image segmentation. However, it often struggles in domains that are either sparsely represented or lie outside its training distribution, such as aerial, medical, and non-RG…

2024

DA-BEV: Unsupervised Domain Adaptation for Bird's Eye View Perception

ECCV 2024poster

"Camera-only Bird’s Eye View (BEV) has demonstrated great potential in environment perception in a 3D space. However, most existing studies were conducted under a supervised setup which cannot scale well while handling various new data. Unsupervised domain adaptive BEV, which effective learning from…

Cited by 5SourcePDFScholar
2024

Domain Adaptation for Large-Vocabulary Object Detectors

NeurIPS 2024poster

Large-vocabulary object detectors (LVDs) aim to detect objects of many categories, which learn super objectness features and can locate objects accurately while applied to various downstream data. However, LVDs often struggle in recognizing the located objects due to domain discrepancy in data distr…

Cited by 1SourcePDFScholar
2024

Historical Test-time Prompt Tuning for Vision Foundation Models

NeurIPS 2024poster

Test-time prompt tuning, which learns prompts online with unlabelled test samples during the inference stage, has demonstrated great potential by learning effective prompts on-the-fly without requiring any task-specific annotations. However, its performance often degrades clearly along the tuning pr…

Cited by 3SourcePDFScholar
2024

MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked Autoencoders

NeurIPS 2024poster

Monocular 3D object detection aims for precise 3D localization and identification of objects from a single-view image. Despite its recent progress, it often struggles while handling pervasive object occlusions that tend to complicate and degrade the prediction of object dimensions, depths, and orien…

Cited by 4SourcePDFScholar
2023

FAC: 3D Representation Learning via Foreground Aware Feature Contrast

CVPR 2023poster

Contrastive learning has recently demonstrated great potential for unsupervised pre-training in 3D scene understanding tasks. However, most existing work randomly selects point features as anchors while building contrast, leading to a clear bias toward background points that often dominate in 3D sce…

2023

Graph Transformer GANs for Graph-Constrained House Generation

CVPR 2023poster

We present a novel graph Transformer generative adversarial network (GTGAN) to learn effective graph node relations in an end-to-end fashion for the challenging graph-constrained house generation task. The proposed graph-Transformer-based generator includes a novel graph Transformer encoder that com…

Cited by 31SourcePDFScholar
2023

High-Resolution Iterative Feedback Network for Camouflaged Object Detection

AAAI 2023technical

Spotting camouflaged objects that are visually assimilated into the background is tricky for both object detection algorithms and humans who are usually confused or cheated by the perfectly intrinsic similarities between the foreground objects and the background surroundings. To tackle this challeng…

2023

Pose-Free Neural Radiance Fields via Implicit Pose Regularization

ICCV 2023poster

Pose-free neural radiance fields (NeRF) aim to train NeRF with unposed multi-view images and it has achieved very impressive success in recent years. Most existing works share the pipeline of training a coarse pose estimator with rendered images at first, followed by a joint optimization of estimate…

Cited by 11PDFScholar
2023

Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic Segmentation

NeurIPS 2023poster

Vision-Language Pre-training has demonstrated its remarkable zero-shot recognition ability and potential to learn generalizable visual representations from languagesupervision. Taking a step ahead, language-supervised semantic segmentation enables spatial localization of textual inputs by learning p…

2023

WaveNeRF: Wavelet-based Generalizable Neural Radiance Fields

ICCV 2023poster

Neural Radiance Field (NeRF) has shown impressive performance in novel view synthesis via implicit scene representation. However, it usually suffers from poor scalability as requiring densely sampled images for each new scene. Several studies have attempted to mitigate this problem by integrating Mu…

Cited by 16PDFScholar
2022

Category Contrast for Unsupervised Domain Adaptation in Visual Tasks

CVPR 2022poster

Instance contrast for unsupervised representation learning has achieved great success in recent years. In this work, we explore the idea of instance contrastive learning in unsupervised domain adaptation (UDA) and propose a novel Category Contrast technique (CaCo) that introduces semantic priors on…

Cited by 199PDFcodeScholar
2022

Generalized Brain Image Synthesis with Transferable Convolutional Sparse Coding Networks

ECCV 2022poster

"High inter-equipment variability and expensive examination costs of brain imaging remain key challenges in leveraging the heterogeneous scans effectively. Despite rapid growth in image-to-image translation with deep learning models, the target brain data may not always be achievable due to the spec…

Cited by 1SourcePDFScholar
2022

GuidedMix-Net: Semi-supervised Semantic Segmentation by Using Labeled Images as Reference

AAAI 2022technical

Semi-supervised learning is a challenging problem which aims to construct a model by learning from limited labeled examples. Numerous methods for this task focus on utilizing the predictions of unlabeled instances consistency alone to regularize networks. However, treating labeled and unlabeled data…

Cited by 26SourcePDFScholar
2022

Hierarchical Variational Memory for Few-shot Learning Across Domains

ICLR 2022poster

Neural memory enables fast adaptation to new tasks with just a few training samples. Existing memory models store features only from the single last layer, which does not generalize well in presence of a domain shift between training and test distributions. Rather than relying on a flat memory, we p…

2022

Highly Accurate Dichotomous Image Segmentation

ECCV 2022poster

"We present a systematic study on a new task called dichotomous image segmentation (DIS), which aims to segment highly accurate objects from natural images. To this end, we collected the first large-scale DIS dataset, called DIS5K, which contains 5,470 high-resolution (e.g., 2K, 4K or larger) images…

2022

Learning Non-Target Knowledge for Few-Shot Semantic Segmentation

CVPR 2022poster

Existing studies in few-shot semantic segmentation only focus on mining the target object information, however, often are hard to tell ambiguous regions, especially in non-target regions, which include background (BG) and Distracting Objects (DOs). To alleviate this problem, we propose a novel frame…

Cited by 152PDFcodeScholar
2022

Learning to Generalize across Domains on Single Test Samples

ICLR 2022poster

We strive to learn a model from a set of source domains that generalizes well to unseen target domains. The main challenge in such a domain generalization scenario is the unavailability of any target domain data during training, resulting in the learned model not being explicitly adapted to the unse…

2022

Multi-Level Representation Learning With Semantic Alignment for Referring Video Object Segmentation

CVPR 2022poster

Referring video object segmentation (RVOS) is a challenging language-guided video grounding task, which requires comprehensively understanding the semantic information of both video content and language queries for object prediction. However, existing methods adopt multi-modal fusion at a frame-base…

Cited by 64PDFScholar
2022

PolarMix: A General Data Augmentation Technique for LiDAR Point Clouds

NeurIPS 2022accept

LiDAR point clouds, which are usually scanned by rotating LiDAR sensors continuously, capture precise geometry of the surrounding environment and are crucial to many autonomous detection and navigation tasks. Though many 3D deep architectures have been developed, efficient collection and annotation…

2022

RePFormer: Refinement Pyramid Transformer for Robust Facial Landmark Detection

IJCAI 2022poster

This paper presents a Refinement Pyramid Transformer (RePFormer) for robust facial landmark detection. Most facial landmark detectors focus on learning representative image features. However, these CNN-based feature representations are not robust enough to handle complex real-world scenarios due to…

Cited by 19SourcePDFScholar
2022

Rethinking Clustering-Based Pseudo-Labeling for Unsupervised Meta-Learning

ECCV 2022poster

"The pioneering unsupervised meta-learning work is a clustering-based pseudo-labeling method, which is model-agnostic and can utilize supervised algorithms for learning from unlabeled data. However, it often suffers from label inconsistency and limited diversity, which leads to poor performance. In…

2022

VITA: A Multi-Source Vicinal Transfer Augmentation Method for Out-of-Distribution Generalization

AAAI 2022technical

Invariance to diverse types of image corruption, such as noise, blurring, or colour shifts, is essential to establish robust models in computer vision. Data augmentation has been the major approach in improving the robustness against common corruptions. However, the samples produced by popular augme…

Cited by 5SourcePDFScholar
2021

A Bit More Bayesian: Domain-Invariant Learning with Uncertainty

ICML 2021spotlight

Domain generalization is challenging due to the domain shift and the uncertainty caused by the inaccessibility of target domain data. In this paper, we address both challenges with a probabilistic framework based on variational Bayesian inference, by incorporating uncertainty into neural network wei…

2021

Architecture Disentanglement for Deep Neural Networks

ICCV 2021poster

Understanding the inner workings of deep neural networks (DNNs) is essential to provide trustworthy artificial intelligence techniques for practical applications. Existing studies typically involve linking semantic concepts to units or layers of DNNs, but fail to explain the inference process. In th…

Cited by 25PDFcodeScholar
2021

Brain Image Synthesis With Unsupervised Multivariate Canonical CSCl4Net

CVPR 2021poster

Recent advances in neuroscience have highlighted the effectiveness of multi-modal medical data for investigating certain pathologies and understanding human cognition. However, obtaining full sets of different modalities is limited by various factors, such as long acquisition times, high examination…

Cited by 8PDFScholar
2021

CCT-Net: Category-Invariant Cross-Domain Transfer for Medical Single-to-Multiple Disease Diagnosis

ICCV 2021poster

A medical imaging model is usually explored for the diagnosis of a single disease. However, with the expanding demand for multi-disease diagnosis in clinical applications, multi-function solutions need to be investigated. Previous works proposed to either exploit different disease labels to conduct…

Cited by 11PDFScholar
2021

D2-Net: Weakly-Supervised Action Localization via Discriminative Embeddings and Denoised Activations

ICCV 2021poster

This work proposes a weakly-supervised temporal action localization framework, called D2-Net, which strives to temporally localize actions using video-level supervision. Our main contribution is the introduction of a novel loss formulation, which jointly enhances the discriminability of latent embed…

Cited by 76PDFcodeScholar
2021

Discriminative Region-Based Multi-Label Zero-Shot Learning

ICCV 2021poster

Multi-label zero-shot learning (ZSL) is a more realistic counter-part of standard single-label ZSL since several objects can co-exist in a natural image. However, the occurrence of multiple objects complicates the reasoning and requires region-specific processing of visual features to preserve their…

Cited by 59PDFcodeScholar
2021

Domain General Face Forgery Detection by Learning to Weight

AAAI 2021technical

In this paper, we propose a domain-general model, termed learning-to-weight (LTW), that guarantees face detection performance across multiple domains, particularly the target domains that are never seen before. However, various face forgery methods cause complex and biased data distributions, making…

2021

Dual-Octave Convolution for Accelerated Parallel MR Image Reconstruction

AAAI 2021technical

Magnetic resonance (MR) image acquisition is an inherently prolonged process, whose acceleration by obtaining multiple undersampled images simultaneously through parallel imaging has always been the subject of research. In this paper, we propose the Dual-Octave Convolution (Dual-OctConv), which is c…

2021

FREE: Feature Refinement for Generalized Zero-Shot Learning

ICCV 2021poster

Generalized zero-shot learning (GZSL) has achieved significant progress, with many efforts dedicated to overcoming the problems of visual-semantic domain gaps and seen-unseen bias. However, most existing methods directly use feature extraction models trained on ImageNet alone, ignoring the cross-dat…

Cited by 244PDFcodeScholar
2021

Full-Duplex Strategy for Video Object Segmentation

ICCV 2021poster

Appearance and motion are two important sources of information in video object segmentation (VOS). Previous methods mainly focus on using simplex solutions, lowering the upper bound of feature collaboration among and across these two cues. In this paper, we study a novel framework, termed the FSNet…

Cited by 180PDFcodeScholar
2021

Generalizable Pedestrian Detection: The Elephant in the Room

CVPR 2021poster

Pedestrian detection is used in many vision based applications ranging from video surveillance to autonomous driving. Despite achieving high performance, it is still largely unknown how well existing detectors generalize to unseen data. This is important because a practical detector should be ready…

Cited by 147PDFcodeScholar
2021

Group Collaborative Learning for Co-Salient Object Detection

CVPR 2021poster

We present a novel group collaborative learning framework (GCNet) capable of detecting co-salient objects in real time (16ms), by simultaneously mining consensus representations at group level based on the two necessary criteria: 1) intra-group compactness to better formulate the consistency among c…

Cited by 115PDFcodeScholar
2021

Group Whitening: Balancing Learning Efficiency and Representational Capacity

CVPR 2021poster

Batch normalization (BN) is an important technique commonly incorporated into deep learning models to perform standardization within mini-batches. The merits of BN in improving a model's learning efficiency can be further amplified by applying whitening, while its drawbacks in estimating population…

Cited by 24PDFcodeScholar
2021

HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot Learning

NeurIPS 2021poster

Zero-shot learning (ZSL) tackles the unseen class recognition problem, transferring semantic knowledge from seen classes to unseen ones. Typically, to guarantee desirable knowledge transfer, a common (latent) space is adopted for associating the visual and semantic domains in ZSL. However, existin…

2021

Kaleido-BERT: Vision-Language Pre-Training on Fashion Domain

CVPR 2021poster

We present a new vision-language (VL) pre-training model dubbed Kaleido-BERT, which introduces a novel kaleido strategy for fashion cross-modality representations from transformers. In contrast to random masking strategy of recent VL models, we design alignment guided masking to jointly focus more o…

Cited by 154PDFcodeScholar
2021

Learning Anchored Unsigned Distance Functions With Gradient Direction Alignment for Single-View Garment Reconstruction

ICCV 2021poster

While single-view 3D reconstruction has made significant progress benefiting from deep shape representations in recent years, garment reconstruction is still not solved well due to open surfaces, diverse topologies and complex geometric details. In this paper, we propose a novel learnable Anchored U…

Cited by 55PDFcodeScholar
2021

Learning To Fuse Asymmetric Feature Maps in Siamese Trackers

CVPR 2021poster

Recently, Siamese-based trackers have achieved promising performance in visual tracking. Most recent Siamese-based trackers typically employ a depth-wise cross-correlation (DW-XCorr) to obtain multi-channel correlation information from the two feature maps (target and search region). However, DW-XCo…

Cited by 97PDFcodeScholar
2021

Light Field Saliency Detection With Dual Local Graph Learning and Reciprocative Guidance

ICCV 2021poster

The application of light field data in salient object detection is becoming increasingly popular in recent years. The difficulty lies in how to effectively fuse the features within the focal stack and how to cooperate them with the feature of the all-focus image. Previous methods usually fuse focal…

Cited by 48PDFcodeScholar
2021

Many-to-One Distribution Learning and K-Nearest Neighbor Smoothing for Thoracic Disease Identification

AAAI 2021technical

Chest X-rays are an important and accessible clinical imaging tool for the detection of many thoracic diseases. Over the past decade, deep learning, with a focus on the convolutional neural network (CNN), has become the most powerful computer-aided diagnosis technology for improving disease identifi…

Cited by 14SourcePDFScholar
2021

MetaNorm: Learning to Normalize Few-Shot Batches Across Domains

ICLR 2021poster

Batch normalization plays a crucial role when training deep neural networks. However, batch statistics become unstable with small batch sizes and are unreliable in the presence of distribution shifts. We propose MetaNorm, a simple yet effective meta-learning normalization. It tackles the aforementio…

Cited by 83SourcePDFScholar
2021

Multi-Stage Progressive Image Restoration

CVPR 2021poster

Image restoration tasks demand a complex balance between spatial details and high-level contextualized information while recovering images. In this paper, we propose a novel synergistic design that can optimally balance these competing goals. Our main proposal is a multi-stage architecture, that pro…

Cited by 2028PDFcodeScholar
2021

Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction Without Convolutions

ICCV 2021poster

Although convolutional neural networks (CNNs) have achieved great success in computer vision, this work investigates a simpler, convolution-free backbone network useful for many dense prediction tasks. Unlike the recently-proposed Vision Transformer (ViT) that was designed for image classification s…

Cited by 5162PDFcodeScholar
2021

RGB-D Saliency Detection via Cascaded Mutual Information Minimization

ICCV 2021poster

Existing RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multi-stage cascaded learning framework via mutual information minimization to explicitly model the multi-modal information between RGB im…

Cited by 139PDFcodeScholar
2021

ReCU: Reviving the Dead Weights in Binary Neural Networks

ICCV 2021poster

Binary neural networks (BNNs) have received increasing attention due to their superior reductions of computation and memory. Most existing works focus on either lessening the quantization error by minimizing the gap between the full-precision weights and their binarization or designing a gradient ap…

Cited by 114PDFcodeScholar
2021

Seminar Learning for Click-Level Weakly Supervised Semantic Segmentation

ICCV 2021poster

Annotation burden has become one of the biggest barriers to semantic segmentation. Approaches based on click-level annotations have therefore attracted increasing attention due to their superior trade-off between supervision and annotation cost. In this paper, we propose seminar learning, a new lear…

Cited by 41PDFScholar
2021

Sparse Needlets for Lighting Estimation With Spherical Transport Loss

ICCV 2021poster

Accurate lighting estimation is challenging yet critical to many computer vision and computer graphics tasks such as high-dynamic-range (HDR) relighting. Existing approaches model lighting in either frequency domain or spatial domain which is insufficient to represent the complex lighting conditions…

Cited by 112PDFScholar
2021

Specificity-Preserving RGB-D Saliency Detection

ICCV 2021poster

RGB-D saliency detection has attracted increasing attention, due to its effectiveness and the fact that depth cues can now be conveniently captured. Existing works often focus on learning a shared representation through various fusion strategies, with few methods explicitly considering how to preser…

Cited by 256PDFcodeScholar
2021

Summarize and Search: Learning Consensus-Aware Dynamic Convolution for Co-Saliency Detection

ICCV 2021poster

Humans perform co-saliency detection by first summarizing the consensus knowledge in the whole group and then searching corresponding objects in each image. Previous methods usually lack robustness, scalability, or stability for the first process and simply fuse consensus features with image feature…

Cited by 72PDFcodeScholar
2021

Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object Segmentation

CVPR 2021poster

This paper addresses the task of unsupervised video multi-object segmentation. Current approaches follow a two-stage paradigm: 1) detect object proposals using pre-trained Mask R-CNN, and 2) conduct generic feature matching for temporal association using re-identification techniques. However, the ge…

Cited by 59PDFScholar
2021

Target-targeted Domain Adaptation for Unsupervised Semantic Segmentation

ICRA 2021poster

Semantic segmentation has attracted increasing attention due to its important role in self-driving, and it is often realized by supervised learning with large number of well labeled maps. However, the labeled images are hard to be obtained in most circumstances, and the common way for unsupervised s…

Cited by 13SourceScholar
2021

TransMatcher: Deep Image Matching Through Transformers for Generalizable Person Re-identification

NeurIPS 2021poster

Transformers have recently gained increasing attention in computer vision. However, existing studies mostly use Transformers for feature representation learning, e.g. for image classification and dense predictions, and the generalizability of Transformers is unknown. In this work, we further investi…

2021

Variational Multi-Task Learning with Gumbel-Softmax Priors

NeurIPS 2021poster

Multi-task learning aims to explore task relatedness to improve individual tasks, which is of particular significance in the challenging scenario that only limited data is available for each task. To tackle this challenge, we propose variational multi-task learning (VMTL), a general probabilistic in…

2021

Visual-Textual Attentive Semantic Consistency for Medical Report Generation

ICCV 2021poster

Diagnosing diseases from medical radiographs and writing reports requires professional knowledge and is time-consuming. To address this, automatic medical report generation approaches have recently gained interest. However, identifying diseases as well as correctly predicting their corresponding siz…

Cited by 24PDFScholar
2020

AnimalWeb: A Large-Scale Hierarchical Dataset of Annotated Animal Faces

CVPR 2020poster

Several studies show that animal needs are often expressed through their faces. Though remarkable progress has been made towards the automatic understanding of human faces, this has not been the case with animal faces. There exists significant room for algorithmic advances that could realize automat…

Cited by 57PDFScholar
2020

BBS-Net: RGB-D Salient Object Detection with a Bifurcated Backbone Strategy Network

ECCV 2020poster

Multi-level feature fusion is a fundamental topic in computer vision for detecting, segmenting, and classifying objects at various scales. When multi-level features meet multi-modal cues, the optimal fusion problem becomes a hot potato. In this paper, we make the first attempt to leverage the inhere…

2020

CLNet: A Compact Latent Network for Fast Adjusting Siamese Trackers

ECCV 2020poster

In this paper, we provide a deep analysis for Siamese-based trackers and find that the one core reason for their failure on challenging cases can be attributed to the problem of {\it decisive samples missing} during offline training. Furthermore, we notice that the samples given in the first frame c…

2020

Count- and Similarity-aware R-CNN for Pedestrian Detection

ECCV 2020poster

Recent pedestrian detection methods generally rely on additional supervision, such as visible bounding-box annotations, to handle heavy occlusions. We propose an approach that leverages pedestrian count and proposal similarity information within a two-stage pedestrian detection framework. Both pedes…

2020

CycleISP: Real Image Restoration via Improved Data Synthesis

CVPR 2020oral

The availability of large-scale datasets has helped unleash the true potential of deep convolutional neural networks (CNNs). However, for the single-image denoising problem, capturing a real dataset is an unacceptably expensive and cumbersome procedure. Consequently, image denoising algorithms are m…

Cited by 450PDFcodeScholar
2020

D2Det: Towards High Quality Object Detection and Instance Segmentation

CVPR 2020poster

We propose a novel two-stage detection method, D2Det, that collectively addresses both precise localization and accurate classification. For precise localization, we introduce a dense local regression that predicts multiple dense box offsets for an object proposal. Different from traditional regress…

Cited by 240PDFcodeScholar
2020

Dynamic Dual-Attentive Aggregation Learning for Visible-Infrared Person Re-Identification

ECCV 2020poster

Visible-infrared person re-identification (VI-ReID) is a challenging cross-modality pedestrian retrieval problem. Due to the large intra-class variations and cross-modality discrepancy with large amount of sample noise, it is difficult to learn discriminative part features. Existing VI-ReID methods…

2020

HRank: Filter Pruning Using High-Rank Feature Map

CVPR 2020oral

Neural network pruning offers a promising prospect to facilitate deploying deep neural networks on resource-limited devices. However, existing methods are still challenged by the training inefficiency and labor cost in pruning designs, due to missing theoretical guidance of non-salient network compo…

Cited by 1040PDFcodeScholar
2020

Hierarchical Human Parsing With Typed Part-Relation Reasoning

CVPR 2020poster

Human parsing is for pixel-wise human semantic understanding. As human bodies are underlying hierarchically structured, how to model human structures is the central theme in this task. Focusing on this, we seek to simultaneously exploit the representational capacity of deep graph networks and the hi…

Cited by 136PDFcodeScholar
2020

Human Parsing Based Texture Transfer from Single Image to 3D Human via Cross-View Consistency

NeurIPS 2020poster

This paper proposes a human parsing based texture transfer model via cross-view consistency learning to generate the texture of 3D human body from a single image. We use the semantic parsing of human body as input for providing both the shape and pose information to reduce the appearance variation…

2020

Interpretable Neural Network Decoupling

ECCV 2020poster

The remarkable performance of convolutional neural networks (CNNs) is entangled with their huge number of uninterpretable parameters, which has become the bottleneck limiting the exploitation of their full potential. Towards network interpretation, previous endeavors mainly resort to the single filt…

Cited by 10SourcePDFScholar
2020

Interpretable and Generalizable Person Re-Identification with Query-Adaptive Convolution and Temporal Lifting

ECCV 2020poster

For person re-identification, existing deep networks often focus on representation learning. However, without transfer learning, the learned model is fixed as is, which is not adaptable for handling various unseen scenarios. In this paper, beyond representation learning, we consider how to formulate…

2020

Latent Embedding Feedback and Discriminative Features for Zero-Shot Classification

ECCV 2020poster

Zero-shot learning strives to classify unseen categories for which no data is available during training. In the generalized variant, the test samples can further belong to seen or unseen categories. The state-of-the-art relies on Generative Adversarial Networks that synthesize unseen class features…

2020

Layer-wise Conditioning Analysis in Exploring the Learning Dynamics of DNNs

ECCV 2020poster

Conditioning analysis uncovers the landscape of an optimization objective by exploring the spectrum of its curvature matrix. This has been well explored theoretically for linear models. We extend this analysis to deep neural networks (DNNs) in order to investigate their learning dynamics. To this en…

Cited by 12SourcePDFScholar
2020

Learning Attentive and Hierarchical Representations for 3D Shape Recognition

ECCV 2020poster

This paper proposes a novel method for 3D shape representation learning, namely Hyperbolic Embedded Attentive Representation (HEAR). Different from existing multi-view based methods, HEAR develops a unified framework to address both multi-view redundancy and single-view incompleteness. Specifically,…

Cited by 36SourcePDFScholar
2020

Learning Enriched Features for Real Image Restoration and Enhancement

ECCV 2020poster

With the goal of recovering high-quality image content from its degraded version, image restoration enjoys numerous applications, such as in surveillance, computational photography and medical imaging. Recently, convolutional neural networks (CNNs) have achieved dramatic improvements over convention…

2020

Learning Multi-Granular Hypergraphs for Video-Based Person Re-Identification

CVPR 2020poster

Video-based person re-identification (re-ID) is an important research topic in computer vision. The key to tackling the challenging task is to exploit both spatial and temporal clues in video sequences. In this work, we propose a novel graph-based framework, namely Multi-Granular Hypergraph (MGH), t…

Cited by 189PDFcodeScholar
2020

Learning to Learn Kernels with Variational Random Features

ICML 2020poster

We introduce kernels with random Fourier features in the meta-learning framework for few-shot learning. We propose meta variational random features (MetaVRF) to learn adaptive kernels for the base-learner, which is developed in a latent variable model by treating the random feature basis as the late…

Cited by 34SourcePDFScholar
2020

Learning to Learn Variational Semantic Memory

NeurIPS 2020poster

In this paper, we introduce variational semantic memory into meta-learning to acquire long-term knowledge for few-shot learning. The variational semantic memory accrues and stores semantic information for the probabilistic inference of class prototypes in a hierarchical Bayesian framework. The seman…

2020

Learning to Learn with Variational Information Bottleneck for Domain Generalization

ECCV 2020poster

Domain generalization models learn to generalize to previously unseen domains, but suffer from prediction uncertainty and domain shift. In this paper, we address both problems. We introduce a probabilistic meta-learning model for domain generalization, in which classifier parameters shared across do…

Cited by 196SourcePDFScholar
2020

Multi-Mutual Consistency Induced Transfer Subspace Learning for Human Motion Segmentation

CVPR 2020poster

Human motion segmentation based on transfer subspace learning is a rising interest in action-related tasks. Although progress has been made, there are still several issues within the existing methods. First, existing methods transfer knowledge from source data to target tasks by learning domain-inva…

Cited by 43PDFScholar
2020

NETNet: Neighbor Erasing and Transferring Network for Better Single Shot Object Detection

CVPR 2020poster

Due to the advantages of real-time detection and improved performance, single-shot detectors have gained great attention recently. To solve the complex scale variations, single-shot detectors make scale-aware predictions based on multiple pyramid layers. However, the features in the pyramid are not…

Cited by 46PDFScholar
2020

On the Number of Linear Regions of Convolutional Neural Networks

ICML 2020poster

One fundamental problem in deep learning is understanding the outstanding performance of deep Neural Networks (NNs) in practice. One explanation for the superiority of NNs is that they can realize a large class of complicated functions, i.e., they have powerful expressivity. The expressivity of a Re…

Cited by 98SourcePDFScholar
2020

Pathological Retinal Region Segmentation From OCT Images Using Geometric Relation Based Augmentation

CVPR 2020poster

Medical image segmentation is important for computer aided diagnosis. Pixelwise manual annotations of large datasets require high expertise and is time consuming. Conventional data augmentations have limited benefit by not fully representing the underlying distribution of the training set, thus affe…

Cited by 46PDFScholar
2020

Region Graph Embedding Network for Zero-Shot Learning

ECCV 2020poster

Most of the existing Zero-Shot Learning (ZSL) approaches learn direct embeddings from global features or image parts (regions) to the semantic space, which, however, fail to capture the appearance relationships between different local regions within a single image. In this paper, to model the relati…

Cited by 195SourcePDFScholar
2020

SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation

ECCV 2020poster

Single-stage instance segmentation approaches have recently gained popularity due to their speed and simplicity, but are still lagging behind in accuracy, compared to two-stage methods. We propose a fast single-stage instance segmentation method, called SipMask, that preserves instance-specific spat…

2020

Super-Resolution and Inpainting with Degraded and Upgraded Generative Adversarial Networks

IJCAI 2020poster

Image super-resolution (SR) and image inpainting are two topical problems in medical image processing. Existing methods for solving the problems are either tailored to recovering a high-resolution version of the low-resolution image or focus on filling missing values, thus inevitably giving rise to…

Cited by 0SourcePDFScholar
2020

Unsupervised Adaptation Learning for Hyperspectral Imagery Super-Resolution

CVPR 2020poster

The key for fusion based hyperspectral image (HSI) super-resolution (SR) is to infer the posteriori of a latent HSI using appropriate image prior and likelihood that depends on degeneration. However, in practice the priors of high-dimensional HSIs can be extremely complicated and the degeneration is…

Cited by 130PDFcodeScholar
2020

Unsupervised Domain Adaptation with Noise Resistible Mutual-Training for Person Re-identification

ECCV 2020poster

Unsupervised domain adaptation (UDA) in the task of person re-identification (re-ID) is highly challenging due to large domain divergence and no class overlap between domains. Pseudo-label based self-training is one of the representative techniques to address UDA. However, label noise caused by unsu…

Cited by 227SourcePDFScholar
2019

3C-Net: Category Count and Center Loss for Weakly-Supervised Action Localization

ICCV 2019poster

Temporal action localization is a challenging computer vision problem with numerous real-world applications. Most existing methods require laborious frame-level supervision to train action localization models. In this work, we propose a framework, called 3C-Net, which only requires video-level super…

Cited by 207PDFcodeScholar
2019

Adversarial Defense by Restricting the Hidden Space of Deep Neural Networks

ICCV 2019poster

Deep neural networks are vulnerable to adversarial attacks which can fool them by adding minuscule perturbations to the input images. The robustness of existing defenses suffers greatly under white-box attack settings, where an adversary has full knowledge about the network and can iterate several t…

Cited by 189PDFcodeScholar
2019

An Iterative and Cooperative Top-Down and Bottom-Up Inference Network for Salient Object Detection

CVPR 2019poster

This paper presents a salient object detection method that integrates both top-down and bottom-up saliency inference in an iterative and cooperative manner. The top-down process is used for coarse-to-fine saliency estimation, where high-level saliency is gradually integrated with finer lower-layer f…

Cited by 259PDFScholar
2019

Attentive Region Embedding Network for Zero-Shot Learning

CVPR 2019poster

Zero-shot learning (ZSL) aims to classify images from unseen categories, by merely utilizing seen class images as the training data. Existing works on ZSL mainly leverage the global features or learn the global regions, from which, to construct the embeddings to the semantic space. However, few of t…

Cited by 351PDFScholar
2019

Building Detail-Sensitive Semantic Segmentation Networks With Polynomial Pooling

CVPR 2019poster

Semantic segmentation is an important computer vision task, which aims to allocate a semantic label to each pixel in an image. When training a segmentation model, it is common to fine-tune a classification network pre-trained on a large-scale dataset. However, as an intrinsic property of the classif…

Cited by 34PDFScholar
2019

Collaborative Learning of Semi-Supervised Segmentation and Classification for Medical Images

CVPR 2019poster

Medical image analysis has two important research areas: disease grading and fine-grained lesion segmentation. Although the former problem often relies on the latter, the two are usually studied separately. Disease severity grading can be treated as a classification problem, which only requires imag…

Cited by 327PDFScholar
2019

Crowd Counting and Density Estimation by Trellis Encoder-Decoder Networks

CVPR 2019poster

Crowd counting has recently attracted increasing interest in computer vision but remains a challenging problem. In this paper, we propose a trellis encoder-decoder network (TEDnet) for crowd counting, which focuses on generating high-quality density estimation maps. The major contributions are four-…

Cited by 435PDFScholar
2019

Deep Contextual Attention for Human-Object Interaction Detection

ICCV 2019poster

Human-object interaction detection is an important and relatively new class of visual relationship detection tasks, essential for deeper scene understanding. Most existing approaches decompose the problem into object localization and interaction recognition. Despite showing progress, these approache…

Cited by 130PDFScholar
2019

Deep Sketch-Shape Hashing With Segmented 3D Stochastic Viewing

CVPR 2019poster

Sketch-based 3D shape retrieval has been extensively studied in recent works, most of which focus on improving the retrieval accuracy, whilst neglecting the efficiency. In this paper, we propose a novel framework for efficient sketch-based 3D shape retrieval, i.e., Deep Sketch-Shape Hashing (DSSH),…

Cited by 48PDFScholar
2019

Efficient Featurized Image Pyramid Network for Single Shot Detector

CVPR 2019poster

Single-stage object detectors have recently gained popularity due to their combined advantage of high detection accuracy and real-time speed. However, while promising results have been achieved by these detectors on standard-sized objects, their performance on small objects is far from satisfactory.…

Cited by 134PDFScholar
2019

Enriched Feature Guided Refinement Network for Object Detection

ICCV 2019poster

We propose a single-stage detection framework that jointly tackles the problem of multi-scale object detection and class imbalance. Rather than designing deeper networks, we introduce a simple yet effective feature enrichment scheme to produce multi-scale contextual features. We further introduce a…

Cited by 109PDFcodeScholar
2019

Gaussian Affinity for Max-Margin Class Imbalanced Learning

ICCV 2019poster

Real-world object classes appear in imbalanced ratios. This poses a significant challenge for classifiers which get biased towards frequent classes. We hypothesize that improving the generalization capability of a classifier should improve learning on imbalanced datasets. Here, we introduce the firs…

Cited by 90PDFScholar
2019

Iterative Normalization: Beyond Standardization Towards Efficient Whitening

CVPR 2019poster

Batch Normalization (BN) is ubiquitously employed for accelerating neural network training and improving the generalization capability by performing standardization within mini-batches. Decorrelated Batch Normalization (DBN) further boosts the above effectiveness by whitening. However, DBN relies…

Cited by 183PDFcodeScholar
2019

Learning Compositional Neural Information Fusion for Human Parsing

ICCV 2019poster

This work proposes to combine neural networks with the compositional hierarchy of human bodies for efficient and complete human parsing. We formulate the approach as a neural information fusion framework. Our model assembles the information from three inference processes over the hierarchy: direct i…

Cited by 160PDFcodeScholar
2019

Learning Rich Features at High-Speed for Single-Shot Object Detection

ICCV 2019poster

Single-stage object detection methods have received significant attention recently due to their characteristic realtime capabilities and high detection accuracies. Generally, most existing single-stage detectors follow two common practices: they employ a network backbone that is pretrained on ImageN…

Cited by 143PDFcodeScholar
2019

Mask-Guided Attention Network for Occluded Pedestrian Detection

ICCV 2019poster

Pedestrian detection relying on deep convolution neural networks has made significant progress. Though promising results have been achieved on standard pedestrians, the performance on heavily occluded pedestrians remains far from satisfactory. The main culprits are intra-class occlusions involving o…

Cited by 252PDFcodeScholar
2019

Object Counting and Instance Segmentation With Image-Level Supervision

CVPR 2019poster

Common object counting in a natural scene is a challenging problem in computer vision with numerous real-world applications. Existing image-level supervised common object counting approaches only predict the global object count and rely on additional instance-level supervision to also determine obje…

Cited by 146PDFcodeScholar
2019

Object-Centric Auto-Encoders and Dummy Anomalies for Abnormal Event Detection in Video

CVPR 2019poster

Abnormal event detection in video is a challenging vision problem. Most existing approaches formulate abnormal event detection as an outlier detection task, due to the scarcity of anomalous data during training. Because of the lack of prior information regarding abnormal events, these methods are no…

Cited by 478PDFScholar
2019

Out-Of-Distribution Detection for Generalized Zero-Shot Action Recognition

CVPR 2019poster

Generalized zero-shot action recognition is a challenging problem, where the task is to recognize new action categories that are unavailable during the training stage, in addition to the seen action categories. Existing approaches suffer from the inherent bias of the learned classifier towards the s…

Cited by 193PDFcodeScholar
2019

RANet: Ranking Attention Network for Fast Video Object Segmentation

ICCV 2019poster

Despite online learning (OL) techniques have boosted the performance of semi-supervised video object segmentation (VOS) methods, the huge time costs of OL greatly restricts their practicality. Matching based and propagation based methods run at a faster speed by avoiding OL techniques. However, they…

Cited by 274PDFcodeScholar
2019

Random Path Selection for Continual Learning

NeurIPS 2019poster

Incremental life-long learning is a main challenge towards the long-standing goal of Artificial General Intelligence. In real-life settings, learning tasks arrive in a sequence and machine learning models must continually learn to increment already acquired knowledge. The existing incremental learni…

2019

Relational Attention Network for Crowd Counting

ICCV 2019poster

Crowd counting is receiving rapidly growing research interests due to its potential application value in numerous real-world scenarios. However, due to various challenges such as occlusion, insufficient resolution and dynamic backgrounds, crowd counting remains an unsolved problem in computer vision…

Cited by 214PDFScholar
2019

See More, Know More: Unsupervised Video Object Segmentation With Co-Attention Siamese Networks

CVPR 2019poster

We introduce a novel network, called as CO-attention Siamese Network (COSNet), to address the unsupervised video object segmentation task from a holistic view. We emphasize the importance of inherent correlation among video frames and incorporate a global co-attention mechanism to improve further th…

Cited by 598PDFcodeScholar
2019

Two Generator Game: Learning to Sample via Linear Goodness-of-Fit Test

NeurIPS 2019poster

Learning the probability distribution of high-dimensional data is a challenging problem. To solve this problem, we formulate a deep energy adversarial network (DEAN), which casts the energy model learned from real data into an optimization of a goodness-of-fit (GOF) test statistic. DEAN can be inter…

Cited by 6SourcePDFScholar
2019

Zero-Shot Video Object Segmentation via Attentive Graph Neural Networks

ICCV 2019oral

This work proposes a novel attentive graph neural network (AGNN) for zero-shot video object segmentation (ZVOS). The suggested AGNN recasts this task as a process of iterative information fusion over video graphs. Specifically, AGNN builds a fully connected graph to efficiently represent frames as n…

Cited by 353PDFcodeScholar
2018

Deep Multi-Task Learning to Recognise Subtle Facial Expressions of Mental States

ECCV 2018poster

Facial expression recognition is a topical task. However, very little research investigates subtle expression recognition, which is important for mental activity analysis, deception detection, etc. We address subtle expression recognition through convolutional neural networks (CNNs) by developing mu…

Cited by 55SourcePDFScholar
2018

Generative Domain-Migration Hashing for Sketch-to-Image Retrieval

ECCV 2018poster

Due to the succinct nature of free-hand sketch drawings, sketch-based image retrieval (SBIR) has abundant practical use cases in consumer electronics. However, SBIR remains a long-standing unsolved problem mainly due to the significant discrepancy between the sketch domain and the image domain. In t…

2018

Highly-Economized Multi-View Binary Compression for Scalable Image Clustering

ECCV 2018poster

How to economically cluster large-scale multi-view images is a long-standing problem in computer vision. To tackle this challenge, this paper introduces a novel approach named Highly-economized Scalable Image Clustering (HSIC) that radically surpasses conventional image clustering methods via binary…

Cited by 55SourcePDFScholar
2018

Hyperparameter Optimization for Tracking With Continuous Deep Q-Learning

CVPR 2018poster

Hyperparameters are numerical presets whose values are assigned prior to the commencement of the learning process. Selecting appropriate hyperparameters is critical for the accuracy of tracking algorithms, yet it is difficult to determine their optimal values, in particular, adaptive ones for each s…

Cited by 198SourcePDFScholar
2018

TBN: Convolutional Neural Network with Ternary Inputs and Binary Weights

ECCV 2018poster

Despite the remarkable success of Convolutional Neural Networks (CNNs) on generalized visual tasks, high computational and memory costs restrict their comprehensive applications on consumer electronics (e.g., portable or smart wearable devices). Recent advancements in binarized networks have demonst…

2018

Towards Universal Representation for Unseen Action Recognition

CVPR 2018poster

Unseen Action Recognition (UAR) aims to recognise novel action categories without training examples. While previous methods focus on inner-dataset seen/unseen splits, this paper proposes a pipeline using a large-scale training source to achieve a Universal Representation (UR) that can generalise to…

Cited by 145SourcePDFScholar
2017

Binary Coding for Partial Action Analysis With Limited Observation Ratios

CVPR 2017poster

Traditional action recognition methods aim to recognize actions with complete observations/executions. However, it is often difficult to capture fully executed actions due to occlusions, interruptions, etc. Meanwhile, action prediction/recognition in advance based on partial observations is essentia…

Cited by 34PDFScholar
2017

Deep Binaries: Encoding Semantic-Rich Cues for Efficient Textual-Visual Cross Retrieval

ICCV 2017poster

Cross-modal hashing is usually regarded as an effective technique for large-scale textual-visual cross retrieval, where data from different modalities are mapped into a shared Hamming space for matching. Most of the traditional textual-visual binary encoding methods only consider holistic image repr…

Cited by 61PDFScholar
2017

Deep Sketch Hashing: Fast Free-Hand Sketch-Based Image Retrieval

CVPR 2017spotlight

Free-hand sketch-based image retrieval (SBIR) is a specific cross-view retrieval task, in which queries are abstract and ambiguous sketches while the retrieval database is formed with natural images. Work in this area mainly focuses on extracting representative and shared features for sketches and n…

Cited by 319PDFcodeScholar
2017

Fast Person Re-Identification via Cross-Camera Semantic Binary Transformation

CVPR 2017poster

Numerous methods have been proposed for person re-identification, most of which however neglect the matching efficiency. Recently, several hashing based approaches have been developed to make re-identification more scalable for large-scale gallery sets. Despite their efficiency, these works ignore c…

Cited by 94PDFScholar
2017

From Zero-Shot Learning to Conventional Supervised Classification: Unseen Visual Data Synthesis

CVPR 2017poster

Robust object recognition systems usually rely on powerful feature extraction mechanisms from a large number of real images. However, in many realistic applications, collecting sufficient images for ever-growing new classes is unattainable. In this paper, we propose a new Zero-shot learning (ZSL) fr…

Cited by 180PDFScholar
2017

Simultaneous Super-Resolution and Cross-Modality Synthesis of 3D Medical Images Using Weakly-Supervised Joint Convolutional Sparse Coding

CVPR 2017poster

Magnetic Resonance Imaging (MRI) offers high-resolution in vivo imaging and rich functional and anatomical multimodality tissue contrast. In practice, however, there are challenges associated with considerations of scanning costs, patient comfort, and scanning time that constrain how much data can b…

Cited by 251PDFScholar
2017

Zero-Shot Action Recognition With Error-Correcting Output Codes

CVPR 2017poster

Recently, zero-shot action recognition (ZSAR) has emerged with the explosive growth of action categories. In this paper, we explore ZSAR from a novel perspective by adopting the Error-Correcting Output Codes (dubbed ZSECOC). Our ZSECOC equips the conventional ECOC with the additional capability of Z…

Cited by 186PDFScholar
2016

Arbitrary view action recognition via transfer dictionary learning on synthetic training data

ICRA 2016

Human action recognition is an important problem in robotic vision. Traditional recognition algorithms usually require the knowledge of view angle, which is not always available in robotic applications such as active vision. In this paper, we propose a new framework to recognize actions with arbitra

Cited by 10SourceScholar