← Search

Xiaoyong Shen

30 accepted papers

2020

Particularity beyond Commonality: Unpaired Identity Transfer with Multiple References

ECCV 2020poster

Unpaired image-to-image translation aims to translate images from the source class to target one by providing sufficient data for these classes. Current few-shot translation methods use multiple reference images to describe the target domain through extracting common features. In this paper, we focu…

Cited by 0SourcePDFScholar
2020

Tensor Low-Rank Reconstruction for Semantic Segmentation

ECCV 2020poster

Context information plays an indispensable role in the success of semantic segmentation. Recently, non-local self-attention based methods are proved to be effective for context information collection. Since desired context consists of spatial-wise and channel-wise attentions, the 3D representation i…

Cited by 88SourcePDFScholar
2019

Associatively Segmenting Instances and Semantics in Point Clouds

CVPR 2019poster

A 3D point cloud describes the real scene precisely and intuitively. To date how to segment diversified elements in such an informative 3D scene is rarely discussed. In this paper, we first introduce a simple and flexible framework to segment instances and semantics in point clouds simultaneously. T…

Cited by 316PDFcodeScholar
2019

Attribute-Driven Spontaneous Motion in Unpaired Image Translation

ICCV 2019poster

Current image translation methods, albeit effective to produce high-quality results in various applications, still do not consider much geometric transform. We in this paper propose the spontaneous motion estimation module, along with a refinement part, to learn attribute-driven deformation between…

Cited by 20PDFcodeScholar
2019

Cross-Domain Adaptation for Animal Pose Estimation

ICCV 2019oral

In this paper, we are interested in pose estimation of animals. Animals usually exhibit a wide range of variations on poses and there is no available animal pose dataset for training and testing. To address this problem, we build an animal pose dataset to facilitate training and evaluation. Consider…

Cited by 223PDFcodeScholar
2019

Dynamic Scene Deblurring With Parameter Selective Sharing and Nested Skip Connections

CVPR 2019poster

Dynamic Scene deblurring is a challenging low-level vision task where spatially variant blur is caused by many factors, e.g., camera shake and object motion. Recent study has made significant progress. Compared with the parameter independence scheme [19] and parameter sharing scheme [33], we develop…

Cited by 414PDFScholar
2019

Fast Point R-CNN

ICCV 2019poster

We present a unified, efficient and effective framework for point-cloud based 3D object detection. Our two-stage approach utilizes both voxel representation and raw point cloud data to exploit respective advantages. The first stage network, with voxel representation as input, only consists of light…

Cited by 511PDFScholar
2019

Hierarchical Point-Edge Interaction Network for Point Cloud Semantic Segmentation

ICCV 2019poster

We achieve 3D semantic scene labeling by exploring semantic relation between each point and its contextual neighbors through edges. Besides an encoder-decoder branch for predicting point labels, we construct an edge branch to hierarchically integrate point features and generate edge features. To inc…

Cited by 245PDFScholar
2019

Learning Shape-Aware Embedding for Scene Text Detection

CVPR 2019poster

We address the problem of detecting scene text in arbitrary shapes, which is a challenging task due to the high variety and complexity of the scene. Specifically, we treat text detection as instance segmentation and propose a segmentation-based framework, which extracts each text instance as an inde…

Cited by 259PDFScholar
2019

MMFace: A Multi-Metric Regression Network for Unconstrained Face Reconstruction

CVPR 2019poster

We propose to address the face reconstruction in the wild by using a multi-metric regression network, MMFace, to align a 3D face morphable model (3DMM) to an input image. The key idea is to utilize a volumetric sub-network to estimate an intermediate geometry representation, and a parametric sub-net…

Cited by 54PDFScholar
2019

Memory-Attended Recurrent Network for Video Captioning

CVPR 2019poster

Typical techniques for video captioning follow the encoder-decoder framework, which can only focus on one source video being processed. A potential disadvantage of such design is that it cannot capture the multiple visual context information of a word appearing in more than one relevant videos in tr…

Cited by 294PDFScholar
2019

Non-Local Recurrent Neural Memory for Supervised Sequence Modeling

ICCV 2019oral

Typical methods for supervised sequence modeling are built upon the recurrent neural networks to capture temporal dependencies. One potential limitation of these methods is that they only model explicitly information interactions between adjacent time steps in a sequence, hence the high-order intera…

Cited by 13PDFcodeScholar
2019

Underexposed Photo Enhancement Using Deep Illumination Estimation

CVPR 2019oral

This paper presents a new neural network for enhancing underexposed photos. Instead of directly learning an image-to-image mapping as previous work, we introduce intermediate illumination in our network to associate the input with expected enhancement result, which augments the network's capability…

Cited by 1084PDFcodeScholar
2018

Facelet-Bank for Fast Portrait Manipulation

CVPR 2018poster

Digital face manipulation has become a popular and fascinating way to touch images with the prevalence of smart phones and social networks. With a wide variety of user preferences, facial expressions, and accessories, a general and flexible model is necessary to accommodate different types of facial…

Cited by 57SourcePDFScholar
2018

ICNet for Real-Time Semantic Segmentation on High-Resolution Images

ECCV 2018poster

We focus on the challenging task of real-time semantic segmentation in this paper. It finds many practical applications and yet is with fundamental difficulty of reducing a large portion of computation for pixel-wise label inference. We propose an image cascade network (ICNet) that incorporates mult…

2018

Image Inpainting via Generative Multi-column Convolutional Neural Networks

NeurIPS 2018poster

In this paper, we propose a generative multi-column network for image inpainting. This network synthesizes different image components in a parallel manner within one stage. To better characterize global structures, we design a confidence-driven reconstruction loss while an implicit diversified MRF r…

2018

Referring Image Segmentation via Recurrent Refinement Networks

CVPR 2018poster

We address the problem of image segmentation from natural language descriptions. Existing deep learning-based methods encode image representations based on the output of the last convolutional layer. One general issue is that the resulting image representation lacks multi-scale semantics, which are…

Cited by 275SourcePDFScholar
2018

Scale-Recurrent Network for Deep Image Deblurring

CVPR 2018poster

In single image deblurring, the ``coarse-to-fine'' scheme, i.e. gradually restoring the sharp image on different resolutions in a pyramid, is very successful in both traditional optimization-based methods and recent neural-network-based approaches. In this paper, we investigate this strategy and pro…

2017

High-Quality Correspondence and Segmentation Estimation for Dual-Lens Smart-Phone Portraits

ICCV 2017poster

Estimating correspondence between two images and extracting the foreground object are two challenges in computer vision. With dual-lens smart phones, such as iPhone 7Plus and Huawei P9, coming into the market, two images of slightly different views provide us new information to unify the two topics.…

Cited by 15PDFScholar
2015

Deep LAC: Deep Localization, Alignment and Classification for Fine-Grained Recognition

CVPR 2015poster

We propose a fine-grained recognition system that incorporates part localization, alignment, and classification in one deep neural network. This is a nontrivial process, as the input to the classification module should be functions that enable back-propagation in constructing the solver. Our major c…

Cited by 439SourcePDFScholar