← Search

Xilin Chen

117 accepted papers

2026

Collaborative Map-Based and Route-Based Policy Learning for Continuous Vision-and-Language Navigation

RA-L 2026

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow language instructions to reach a target in unseen, 3D environments. A powerful VLN-CE agent requires two crucial abilities during cross-modal planning: spatial reasoning to explore towards the target locat

Cited by 0SourceScholar
2026

Contrastive Spectral Rectification: Test-Time Defense towards Zero-shot Adversarial Robustness of CLIP

ICML 2026poster

Vision-language models (VLMs) such as CLIP have demonstrated remarkable zero-shot generalization, yet remain highly vulnerable to adversarial examples (AEs). While test-time defenses are promising, existing methods fail to provide sufficient robustness against strong attacks and are often hampered b…

Cited by 0SourceScholar
2026

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes

ICLR 2026poster

The aspiration for artificial general intelligence, fueled by the rapid progress of multimodal understanding, demands models to understand humans in diverse and complex scenarios, as humans manifests intelligence and embody the world. We propose HumanPCR, an evaluation suite for probing MLLMs’ capac…

Cited by 0SourceScholar
2026

MAST: Motif-Augmented Diffusion with Search Tree for Spectroscopic Molecular Structure Elucidation

ICML 2026poster

Elucidating molecular structures from spectra is a foundational problem in chemical and materials characterization, yet remains challenging due to spectral ambiguity and the vast molecular space. Although recent diffusion-based generators show strong promise for spectra-conditioned elucidation, exis…

Cited by 0SourceScholar
2026

Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling

ICLR 2026poster

Safe and feasible trajectory planning is critical for real-world autonomous driving systems. However, existing learning-based planners rely heavily on expert demonstrations, which not only lack explicit safety awareness but also risk inheriting undesirable behaviors such as speeding from suboptimal…

Cited by 0SourcecodeScholar
2026

RoboPCA: Pose-Centered Affordance Learning from Human Demonstrations for Robot Manipulation

ICRA 2026poster

Understanding spatial affordances---comprising the contact regions of object interaction and the corresponding contact poses---is essential for robots to effectively manipulate objects and accomplish diverse tasks. However, existing spatial affordance prediction methods mainly focus on locating the …

2026

SkillNet: Hierarchical Skill Modeling for Compositional Generalization in Vision-Language Action Models

ICML 2026poster

Transfer across diverse task compositions and unseen behaviors remains a significant challenge for vision-language action (VLA) models. Skills are repeatable and atomic components for various tasks, and similarities shared with different skills provide evidence for transferability across behaviors. …

Cited by 0SourceScholar
2026

Tell as You Want: Customizing Image Narrative with Knowledge and Thoughts

AAAI 2026technical

With the advancement of vision-language models, image captioning has made significant progress, leading to the generation of more accurate and detailed descriptions. Current image captioning primarily focuses on describing the apparent visual characteristics, which are easily observed by most humans

Cited by 0SourcePDFScholar
2026

Uni-Skill: Building Self-Evolving Skill Repository for Generalizable Robotic Manipulation

ICRA 2026poster

While skill-centric approaches leverage foundation models to enhance generalization in compositional tasks, they often rely on fixed skill libraries, limiting adaptability to new tasks without manual intervention. To address this, we propose Uni-Skill, a Unified Skill-centric framework that supports…

2026

UniPercept: A Unified Diffusion Model for Generalizable Visual Perception

CVPR 2026

Diffusion models have shown impressive performance in generative tasks, demonstrating their ability to capture detailed structural and semantic information. Recently, these capabilities have been extended to visual understanding, with studies employing diffusion models as the backbone for various pe

Cited by 0SourcecodeScholar
2026

V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs

CVPR 2026

Adversarial attacks have evolved from simply disrupting predictions on conventional task-specific models to the more complex goal of manipulating image semantics in Large Vision-Language Models (LVLMs). However, existing methods struggle with controllability and cannot precisely manipulate the seman

Cited by 0SourcecodeScholar
2026

Walking World Model for Visually Impaired Path Following

RA-L 2026

Guiding visually impaired individuals (VI) walking along planned paths is essential for enabling independent long-distance mobility. Current reactive approaches only correct deviations after they occur. These methods ignore VI's walking dynamics (e.g., reaction latency and heading drift), resulting

Cited by 0SourceScholar
2025

CogCM: Cognition-Inspired Contextual Modeling for Audio-Visual Speech Enhancement

ICCV 2025poster

Audio-Visual Speech Enhancement (AVSE) leverages both audio and visual information to improve speech quality. Despite noisy real-world conditions, humans are generally able to perceive and interpret corrupted speech segments as clear. Researches in cognitive science have shown how the brain merges a…

Cited by 0SourcePDFScholar
2025

CtrLoRA: An Extensible and Efficient Framework for Controllable Image Generation

ICLR 2025poster

Recently, large-scale diffusion models have made impressive progress in text-to-image (T2I) generation. To further equip these T2I models with fine-grained spatial control, approaches like ControlNet introduce an extra network that learns to follow a condition image. However, for every single condit…

2025

Dynamic Behavior Cloning With Temporal Feature Prediction: Enhancing Robotic Arm Manipulation in Moving Object Tasks

RA-L 2025

In numerous real-world applications, the ability to accurately perceive and respond to dynamic changes in the environment, while also maintaining the flexibility to transfer learned skills across different tasks, is crucial for the effective operation of robotic arms. Behavior cloning is particularl

Cited by 5SourceScholar
2025

Dysca: A Dynamic and Scalable Benchmark for Evaluating Perception Ability of LVLMs

ICLR 2025poster

Currently many benchmarks have been proposed to evaluate the perception ability of the Large Vision-Language Models (LVLMs). However, most benchmarks conduct questions by selecting images from existing datasets, resulting in the potential data leakage. Besides, these benchmarks merely focus on evalu…

2025

EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models

ICCV 2025poster

The progress on generative models has led to significant advances on text-to-video (T2V) generation, yet the motion controllability of generated videos remains limited. Existing motion transfer approaches explored the motion representations of reference videos to guide generation. Nevertheless, thes…

2025

G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction via Evolutionary Diffusion

ICCV 2025poster

Understanding how genes influence phenotype across species is a fundamental challenge in genetic engineering, which will facilitate advances in various fields such as crop breeding, conservation biology, and personalized medicine. However, current phenotype prediction models are limited to individua…

Cited by 0SourcePDFScholar
2025

KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge

NeurIPS 2025poster

The molecular large language models have garnered widespread attention due to their promising potential on molecular applications. However, current molecular large language models face significant limitations in understanding molecules due to inadequate textual descriptions and suboptimal molecular…

Cited by 0SourceScholar
2025

MATS: An Audio Language Model under Text-only Supervision

ICML 2025poster

Large audio-language models (LALMs), built upon powerful Large Language Models (LLMs), have exhibited remarkable audio comprehension and reasoning capabilities. However, the training of LALMs demands a large corpus of audio-language pairs, which requires substantial costs in both data collection an…

2025

Not Only Vision: Evolve Visual Speech Recognition via Peripheral Information

ICCV 2025poster

Is visual information alone sufficient for visual speech recognition (VSR) in challenging real-world scenarios? Humans do not rely solely on visual information for lip-reading but also incorporate additional cues, such as speech-related context and prior knowledge about the task. However, existing m…

Cited by 0SourcePDFScholar
2025

OV3D-CG: Open-vocabulary 3D Instance Segmentation with Contextual Guidance

ICCV 2025poster

Open-vocabulary 3D instance segmentation (OV-3DIS), which aims to segment and classify objects beyond predefined categories, is a critical capability for embodied AI applications. Existing methods rely on pre-trained 2D foundation models, focusing on instance-level features while overlooking context…

2025

ProtInvTree: Deliberate Protein Inverse Folding with Reward-guided Tree Search

NeurIPS 2025spotlight

Designing protein sequences that fold into a target 3D structure—known as protein inverse folding—is a fundamental challenge in protein engineering. While recent deep learning methods have achieved impressive performance by recovering native sequences, they often overlook the one-to-many nature of t…

Cited by 0SourcecodeScholar
2025

R2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action Planner

CVPR 2025poster

This paper explores using large language models (LLMs) as low-level action planners for embodied tasks. While LLMs excel as the robot's "brain" for high-level planning, they face challenges in directly controlling the "body" by generating precise low-level actions. This limitation arises from LLMs'…

2025

Revisiting Logit Distributions for Reliable Out-of-Distribution Detection

NeurIPS 2025poster

Out-of-distribution (OOD) detection is critical for ensuring the reliability of deep learning models in open-world applications. While post-hoc methods are favored for their efficiency and ease of deployment, existing approaches often underexploit the rich information embedded in the model’s logits…

Cited by 0SourcecodeScholar
2025

Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation

IROS 2025

Zero-shot generalization across various robots, tasks and environments remains a significant challenge in robotic manipulation. Policy code generation methods use executable code to connect high-level task descriptions and low-level action sequences, leveraging the generalization capabilities of lar

Cited by 5SourcecodeScholar
2025

UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing

CVPR 2025highlight

Human pose plays a crucial role in the digital age. While recent works have achieved impressive progress in understanding and generating human poses, they often support only a single modality of control signals and operate in isolation, limiting their application in real-world scenarios. This paper…

2025

Wavelet-Driven Masked Image Modeling: A Path to Efficient Visual Representation

AAAI 2025technical

Masked Image Modeling (MIM) has garnered significant attention in self-supervised learning, thanks to its impressive capacity to learn scalable visual representations tailored for downstream tasks. However, images inherently contain abundant redundant information, leading the pixel-based MIM reconst…

Cited by 0SourcePDFScholar
2025

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP

NeurIPS 2025poster

Contrastive Language-Image Pre-training (CLIP) has become a foundation model and has been applied to various vision and multimodal tasks. However, recent works indicate that CLIP falls short in distinguishing detailed differences in images and shows suboptimal performance on dense-prediction and vis…

Cited by 0SourcecodeScholar
2024

A Simple Romance Between Multi-Exit Vision Transformer and Token Reduction

ICLR 2024poster

Vision Transformers (ViTs) are now flourishing in the computer vision area. Despite the remarkable success, ViTs suffer from high computational costs, which greatly hinder their practical usage. Token reduction, which identifies and discards unimportant tokens during forward propagation, has then be…

Cited by 8SourcePDFScholar
2024

An Information Theoretical View for Out-Of-Distribution Detection

ECCV 2024poster

"Detecting out-of-distribution (OOD) inputs are pivotal for real-world applications. However, due to the inaccessibility of OODs during training phase, applying supervised binary classification with in-distribution (ID) and OOD labels is not feasible. Therefore, previous works typically employ the p…

Cited by 0SourcePDFScholar
2024

ES3: Evolving Self-Supervised Learning of Robust Audio-Visual Speech Representations

CVPR 2024poster

We propose a novel strategy ES3 for self-supervised learning of robust audio-visual speech representations from unlabeled talking face videos. While many recent approaches for this task primarily rely on guiding the learning process using the audio modality alone to capture information shared betwee…

Cited by 2SourcePDFScholar
2024

HPNet: Dynamic Trajectory Forecasting with Historical Prediction Attention

CVPR 2024poster

Predicting the trajectories of road agents is essential for autonomous driving systems. The recent mainstream methods follow a static paradigm which predicts the future trajectory by using a fixed duration of historical frames. These methods make the predictions independently even at adjacent time s…

2024

HiFi-Score: Fine-grained Image Description Evaluation with Hierarchical Parsing Graphs

ECCV 2024poster

"With the advancements of vision-language models, the growing demand for generating customized image descriptions under length, target regions, and other various control conditions brings new challenges for evaluation. Most existing metrics, designed primarily for single-sentence image captioning wi…

2024

Point2Real: Bridging the Gap between Point Cloud and Realistic Image for Open-World 3D Recognition

AAAI 2024technical

Recognition in open-world scenarios is an important and challenging field, where Vision-Language Pre-training paradigms have greatly impacted the 2D domain. This inspires a growing interest in introducing 2D pre-trained models, such as CLIP, into the 3D domain to enhance the ability of point cloud u…

2024

Rethinking the Evaluation of Out-of-Distribution Detection: A Sorites Paradox

NeurIPS 2024poster

Most existing out-of-distribution (OOD) detection benchmarks classify samples with novel labels as the OOD data. However, some marginal OOD samples actually have close semantic contents to the in-distribution (ID) sample, which makes determining the OOD sample a Sorites Paradox. In this paper, we co…

2024

Scalable Modular Network: A Framework for Adaptive Learning via Agreement Routing

ICLR 2024poster

In this paper, we propose a novel modular network framework, called Scalable Modular Network (SMN), which enables adaptive learning capability and supports integration of new modules after pre-training for better adaptation. This adaptive capability comes from a novel design of router within SMN, na…

2024

T2IShield: Defending Against Backdoors on Text-to-Image Diffusion Models

ECCV 2024poster

"While text-to-image diffusion models demonstrate impressive generation capabilities, they also exhibit vulnerability to backdoor attacks, which involve the manipulation of model outputs through malicious triggers. In this paper, for the first time, we propose a comprehensive defense method named T2…

2024

Think before Placement: Common Sense Enhanced Transformer for Object Placement

ECCV 2024poster

"Object placement is a task to insert a foreground object into a background scene at a suitable position and size. Existing methods mainly focus on extracting better visual features, while neglecting common sense about the objects and background. It leads to semantically unrealistic object positions…

2024

UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language Models

NeurIPS 2024poster

Pre-trained vision-language models (e.g., CLIP) have shown powerful zero-shot transfer capabilities. But they still struggle with domain shifts and typically require labeled data to adapt to downstream tasks, which could be costly. In this work, we aim to leverage unlabeled data that naturally span…

2023

CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language Recognition

ICCV 2023poster

The co-occurrence signals (e.g., hand shape, facial expression, and lip pattern) play a critical role in Continuous Sign Language Recognition (CSLR). Compared to RGB data, skeleton data provide a more efficient and concise option, and lay a good foundation for the co-occurrence exploration in CSLR.…

Cited by 27PDFScholar
2023

DISC: Learning From Noisy Labels via Dynamic Instance-Specific Selection and Correction

CVPR 2023poster

Existing studies indicate that deep neural networks (DNNs) can eventually memorize the label noise. We observe that the memorization strength of DNNs towards each instance is different and can be represented by the confidence value, which becomes larger and larger during the training process. Based…

2023

DandelionNet: Domain Composition with Instance Adaptive Classification for Domain Generalization

ICCV 2023poster

Domain generalization (DG) attempts to learn a model on source domains that can well generalize to unseen but different domains. The multiple source domains are innately different in distribution but intrinsically related to each other, e.g., from the same label space. To achieve a generalizable fea…

Cited by 7PDFScholar
2023

Generalized Semi-Supervised Learning via Self-Supervised Feature Adaptation

NeurIPS 2023poster

Traditional semi-supervised learning (SSL) assumes that the feature distributions of labeled and unlabeled data are consistent which rarely holds in realistic scenarios. In this paper, we propose a novel SSL setting, where unlabeled samples are drawn from a mixed distribution that deviates from the…

Cited by 5SourcePDFScholar
2023

Glance and Focus: Memory Prompting for Multi-Event Video Question Answering

NeurIPS 2023poster

Video Question Answering (VideoQA) has emerged as a vital tool to evaluate agents’ ability to understand human daily behaviors. Despite the recent success of large vision language models in many multi-modal tasks, complex situation reasoning over videos involving multiple human-object interaction ev…

2023

Source-Free Adaptive Gaze Estimation by Uncertainty Reduction

CVPR 2023poster

Gaze estimation across domains has been explored recently because the training data are usually collected under controlled conditions while the trained gaze estimators are used in real and diverse environments. However, due to privacy and efficiency concerns, simultaneous access to annotated source…

2023

Understanding Few-Shot Learning: Measuring Task Relatedness and Adaptation Difficulty via Attributes

NeurIPS 2023poster

Few-shot learning (FSL) aims to learn novel tasks with very few labeled samples by leveraging experience from \emph{related} training tasks. In this paper, we try to understand FSL by exploring two key questions: (1) How to quantify the relationship between \emph{ training} and \emph{novel}…

2022

Clothes-Changing Person Re-Identification With RGB Modality Only

CVPR 2022poster

The key to address clothes-changing person re-identification (re-id) is to extract clothes-irrelevant features, e.g., face, hairstyle, body shape, and gait. Most current works mainly focus on modeling body shape from multi-modality information (e.g., silhouettes and sketches), but do not make full u…

Cited by 225PDFcodeScholar
2022

Deep Radial Embedding for Visual Sequence Learning

ECCV 2022poster

"Connectionist Temporal Classification (CTC) is a popular objective function in sequence recognition, which provides supervision for unsegmented sequence data through aligning sequence and its corresponding labeling iteratively. The blank class of CTC plays a crucial role in the alignment process an…

Cited by 20SourcePDFScholar
2022

Implicit-Part Based Context Aggregation for Point Cloud Instance Segmentation

IROS 2022poster

Context information is important for instance segmentation on point clouds. Existing methods either only use local surroundings by stacking multiple convolution layers or use non-local methods to model long-range interactions. However, they usually directly operate on points which is an unstructured…

Cited by 0SourcecodeScholar
2022

Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework

ECCV 2022poster

"The current popular two-stream, two-stage tracking framework extracts the template and the search region features separately and then performs relation modeling, thus the extracted features lack the awareness of the target and have limited target-background discriminability. To tackle the above iss…

2022

Optimal Positive Generation via Latent Transformation for Contrastive Learning

NeurIPS 2022accept

Contrastive learning, which learns to contrast positive with negative pairs of samples, has been popular for self-supervised visual representation learning. Although great effort has been made to design proper positive pairs through data augmentation, few works attempt to generate optimal positives…

Cited by 10SourcePDFScholar
2022

Salient-to-Broad Transition for Video Person Re-Identification

CVPR 2022poster

Due to the limited utilization of temporal relations in video re-id, the frame-level attention regions of mainstream methods are partial and highly similar. To address this problem, we propose a Salient-to-Broad Module (SBM) to enlarge the attention regions gradually. Specifically, in SBM, while the…

Cited by 69PDFcodeScholar
2021

Env-QA: A Video Question Answering Benchmark for Comprehensive Understanding of Dynamic Environments

ICCV 2021poster

Visual understanding goes well beyond the study of images or videos on the web. To achieve complex tasks in volatile situations, the human can deeply understand the environment, quickly perceive events happening around, and continuously track objects' state changes, which are still challenging for c…

Cited by 37PDFScholar
2021

FAIEr: Fidelity and Adequacy Ensured Image Caption Evaluation

CVPR 2021poster

Image caption evaluation is a crucial task, which involves the semantic perception and matching of image and text. Good evaluation metrics aim to be fair, comprehensive, and consistent with human judge intentions. When humans evaluate a caption, they usually consider multiple aspects, such as whethe…

Cited by 41PDFcodeScholar
2021

HRFormer: High-Resolution Vision Transformer for Dense Predict

NeurIPS 2021poster

We present a High-Resolution Transformer (HRFormer) that learns high-resolution representations for dense prediction tasks, in contrast to the original Vision Transformer that produces low-resolution representations and has high memory and computational cost. We take advantage of the multi-resolutio…

2021

Hierarchical Context-aware Network for Dense Video Event Captioning

ACL 2021long

Dense video event captioning aims to generate a sequence of descriptive captions for each event in a long untrimmed video. Video-level context provides important information and facilities the model to generate consistent and less redundant captions between events. In this paper, we introduce a nove…

2021

Holistic Pose Graph: Modeling Geometric Structure Among Objects in a Scene Using Graph Inference for 3D Object Prediction

ICCV 2021poster

Due to the missing depth cues, it is essentially ambiguous to detect 3D objects from a single RGB image. Existing methods predict the 3D pose for each object independently or merely by combining local relationships within limited surroundings, but rarely explore the inherent geometric relationships…

Cited by 5PDFScholar
2021

Visual Alignment Constraint for Continuous Sign Language Recognition

ICCV 2021poster

Vision-based Continuous Sign Language Recognition (CSLR) aims to recognize unsegmented signs from image streams. Overfitting is one of the most critical problems in CSLR training, and previous works show that the iterative training scheme can partially solve this problem while also costing more trai…

Cited by 189PDFcodeScholar
2020

An Efficient PointLSTM for Point Clouds Based Gesture Recognition

CVPR 2020poster

Point clouds contain rich spatial information, which provides complementary cues for gesture recognition. In this paper, we formulate gesture recognition as an irregular sequence recognition problem and aim to capture long-term spatial correlations across point cloud sequences. A novel and effective…

Cited by 136PDFcodeScholar
2020

Appearance-Preserving 3D Convolution for Video-based Person Re-identification

ECCV 2020poster

Due to the imperfect person detection results and posture changes, temporal appearance misalignment is unavoidable in video-based person re-identification (ReID). In this case, 3D convolution may destroy the appearance representation of person video clips, thus it is harmful to ReID. To address this…

2020

Cross-Domain Face Presentation Attack Detection via Multi-Domain Disentangled Representation Learning

CVPR 2020poster

Face presentation attack detection (PAD) has been an urgent problem to be solved in the face recognition systems. Conventional approaches usually assume the testing and training are within the same domain; as a result, they may not generalize well into unseen scenarios because the representations le…

Cited by 234PDFcodeScholar
2020

Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training

ECCV 2020poster

Although two-stage object detectors have continuously advanced the state-of-the-art performance in recent years, the training process itself is far from crystal. In this work, we first point out the inconsistency problem between the fixed network settings and the dynamic training procedure, which gr…

2020

Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene Text

CVPR 2020poster

Answering questions that require reading texts in an image is challenging for current models. One key difficulty of this task is that rare, polysemous, and ambiguous words frequently appear in images, e.g., names of places, products, and sports teams. To overcome this difficulty, only resorting to p…

Cited by 150PDFScholar
2020

SegFix: Model-Agnostic Boundary Refinement for Segmentation

ECCV 2020poster

We present a model-agnostic post-processing scheme to improve the boundary quality for the segmentation result that is generated by any existing segmentation model. Motivated by the empirical observation that the label predictions of interior pixels are more reliable, we propose to replace the origi…

2020

Self-Supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation

CVPR 2020oral

Image-level weakly supervised semantic segmentation is a challenging problem that has been deeply studied in recent years. Most of advanced solutions exploit class activation map (CAM). However, CAMs can hardly serve as the object mask due to the gap between full and weak supervisions. In this paper…

Cited by 881PDFcodeScholar
2020

Sketching Image Gist: Human-Mimetic Hierarchical Scene Graph Generation

ECCV 2020poster

Scene graph aims to faithfully reveal humans' perception of image content. When humans analyze a scene, they usually prefer to describe image gist first, namely major objects and key relations in a scene graph. This humans' inherent perceptive habit implies that there exists a hierarchical structure…

2020

TCTS: A Task-Consistent Two-Stage Framework for Person Search

CVPR 2020poster

The state of the art person search methods separate person search into detection and re-ID stages, but ignore the consistency between these two stages. The general person detector has no special attention on the query target; The re-ID model is trained on hand-drawn bounding boxes which are not avai…

Cited by 139PDFScholar
2020

Temporal Complementary Learning for Video Person Re-Identification

ECCV 2020poster

This paper proposes a Temporal Complementary Learning Network that extracts complementary features of consecutive video frames for video person re-identification. Firstly, we introduce a Temporal Saliency Erasing (TSE) module including a saliency erasing operation and a series of ordered learners. S…

2020

Unsupervised Domain Adaptation With Hierarchical Gradient Synchronization

CVPR 2020poster

Domain adaptation attempts to boost the performance on a target domain by borrowing knowledge from a well established source domain. To handle the distribution gap between two domains, the prominent approaches endeavor to extract domain-invariant features. It is known that after a perfect domain ali…

Cited by 121PDFScholar
2019

Cross Attention Network for Few-shot Classification

NeurIPS 2019poster

Few-shot classification aims to recognize unlabeled samples from unseen classes given only few labeled samples. The unseen classes and low-data problem make few-shot classification very challenging. Many existing approaches extracted features from labeled and unlabeled samples independently, as a re…

2019

Exploring Context and Visual Pattern of Relationship for Scene Graph Generation

CVPR 2019poster

Relationship is the core of scene graph, but its prediction is far from satisfying because of its complex visual diversity. To alleviate this problem, we treat relationship as an abstract object, exploring not only significative visual pattern but contextual information for it, which are two key asp…

Cited by 115PDFScholar
2019

Fully Learnable Group Convolution for Acceleration of Deep Neural Networks

CVPR 2019poster

Benefitted from its great success on many tasks, deep learning is increasingly used on low-computational-cost devices, e.g. smartphone, embedded devices, etc. To reduce the high computational and memory cost, in this work, we propose a fully learnable group convolution module (FLGC for short) which…

Cited by 94PDFScholar
2019

Interaction-And-Aggregation Network for Person Re-Identification

CVPR 2019poster

Person re-identification (reID) benefits greatly from deep convolutional neural networks (CNNs) which learn robust feature embeddings. However, CNNs are inherently limited in modeling the large variations in person pose and scale due to their fixed geometric structures. In this paper, we propose a n…

Cited by 466PDFScholar
2019

Multi-label Co-regularization for Semi-supervised Facial Action Unit Recognition

NeurIPS 2019poster

Facial action units (AUs) recognition is essential for emotion analysis and has been widely applied in mental state analysis. Existing work on AU recognition usually requires big face dataset with accurate AU labels. However, manual AU annotation requires expertise and can be time-consuming. In this…

2019

S2GAN: Share Aging Factors Across Ages and Share Aging Trends Among Individuals

ICCV 2019oral

Generally, we human follow the roughly common aging trends, e.g., the wrinkles only tend to be more, longer or deeper. However, the aging process of each individual is more dominated by his/her personalized factors, including the invariant factors such as identity and mole, as well as the personaliz…

Cited by 57PDFScholar
2019

Self-Supervised Representation Learning From Videos for Facial Action Unit Detection

CVPR 2019oral

In this paper, we aim to learn discriminative representation for facial action unit (AU) detection from large amount of videos without manual annotations. Inspired by the fact that facial actions are the movements of facial muscles, we depict the movements as the transformation between two face imag…

Cited by 129PDFcodeScholar
2019

Temporal Knowledge Propagation for Image-to-Video Person Re-Identification

ICCV 2019poster

In many scenarios of Person Re-identification (Re-ID), the gallery set consists of lots of surveillance videos and the query is just an image, thus Re-ID has to be conducted between image and videos. Compared with videos, still person images lack temporal information. Besides, the information asymme…

Cited by 85PDFcodeScholar
2019

Transferable Contrastive Network for Generalized Zero-Shot Learning

ICCV 2019poster

Zero-shot learning (ZSL) is a challenging problem that aims to recognize the target categories without seen data, where semantic information is leveraged to transfer knowledge from some source classes. Although ZSL has made great progress in recent years, most existing approaches are easy to overfit…

Cited by 241PDFScholar
2019

VRSTC: Occlusion-Free Video Person Re-Identification

CVPR 2019poster

Video person re-identification (re-ID) plays an important role in surveillance video analysis. However, the performance of video re-ID degenerates severely under partial occlusion. In this paper, we propose a novel network, called Spatio-Temporal Completion network (STCnet), to explicitly handle par…

Cited by 271PDFScholar
2018

Duplex Generative Adversarial Network for Unsupervised Domain Adaptation

CVPR 2018poster

Domain adaptation attempts to transfer the knowledge obtained from the source domain to the target domain, i.e., the domain where the testing data are. The main challenge lies in the distribution discrepancy between source and target domain. Most existing works endeavor to learn domain invariant rep…

Cited by 216SourcePDFScholar
2018

Facial Expression Recognition with Inconsistently Annotated Datasets

ECCV 2018poster

Annotation errors and bias are inevitable among different facial expression datasets due to the subjectiveness of annotating facial expressions. Ascribe to the inconsistent annotations, performance of existing facial expression recognition (FER) methods cannot keep improving when the training set is…

2018

Generative Adversarial Network with Spatial Attention for Face Attribute Editing

ECCV 2018poster

Face attribute editing aims at editing the face image with the given attribute. Most existing works employ Generative Adversarial Network (GAN) to operate face attribute editing. However, these methods inevitably change the attribute-irrelevant regions, as shown in Fig.~ ef{fig1}. Therefore, we intr…

Cited by 189SourcePDFScholar
2018

Learning Class Prototypes via Structure Alignment for Zero-Shot Recognition

ECCV 2018poster

Zero-shot learning (ZSL) aims to recognize objects of novel classes without any training samples of specific classes, which is achieved by exploiting the semantic information and auxiliary datasets. Recently most ZSL approaches focus on learning visual-semantic embeddings to transfer knowledge from…

Cited by 162SourcePDFScholar
2018

Real-Time Rotation-Invariant Face Detection With Progressive Calibration Networks

CVPR 2018poster

Rotation-invariant face detection, i.e. detecting faces with arbitrary rotation-in-plane (RIP) angles, is widely required in unconstrained applications but still remains as a challenging task, due to the large variations of face appearances. Most existing methods compromise with speed or accuracy to…

2018

Structure Inference Net: Object Detection Using Scene-Level Context and Instance-Level Relationships

CVPR 2018poster

Context is important for accurate visual recognition. In this work we propose an object detection algorithm that not only considers object visual appearance, but also makes use of two kinds of context including scene contextual information and object relationships within a single image. Therefore, o…

Cited by 298SourcePDFScholar
2017

Discriminative Covariance Oriented Representation Learning for Face Recognition With Image Sets

CVPR 2017poster

For face recognition with image sets, while most existing works mainly focus on building robust set models with hand-crafted feature, it remains a research gap to learn better image representations which can closely match the subsequent image set modeling and classification. Taking sample covariance…

Cited by 49PDFScholar
2017

Learning Discriminative Latent Attributes for Zero-Shot Classification

ICCV 2017poster

Zero-shot learning (ZSL) aims to transfer knowledge from observed classes to the unseen classes, based on the assumption that both the seen and unseen classes share a common semantic space, among which attributes enjoy a great popularity. However, few works study whether the human-designed semantic…

Cited by 128PDFScholar
2017

Learning Multifunctional Binary Codes for Both Category and Attribute Oriented Retrieval Tasks

CVPR 2017poster

In this paper we propose a unified framework to address multiple realistic image retrieval tasks concerning both category and attributes. Considering the scale of modern datasets, hashing is favorable for its low complexity. However, most existing hashing methods are designed to preserve one single…

Cited by 41PDFScholar
2017

Recursive Spatial Transformer (ReST) for Alignment-Free Face Recognition

ICCV 2017spotlight

Convolutional Neural Network (CNN) has led to significant progress in face recognition. Currently most CNN-based face recognition methods follow a two-step pipeline, i.e. a detected face is first aligned to a canonical one pre-defined by a mean face shape, and then it is fed into a CNN to extract fe…

Cited by 59PDFScholar
2016

Occlusion-Free Face Alignment: Deep Regression Networks Coupled With De-Corrupt AutoEncoders

CVPR 2016poster

Face alignment or facial landmark detection plays an important role in many computer vision applications, e.g., face recognition, facial expression recognition, face animation, etc. However, the performance of face alignment system degenerates severely when occlusions occur. In this work, we propose…

Cited by 128PDFScholar
2015

Discriminant Analysis on Riemannian Manifold of Gaussian Distributions for Face Recognition With Image Sets

CVPR 2015poster

This paper presents a method named Discriminant Analysis on Riemannian manifold of Gaussian distributions (DARG) to solve the problem of face recognition with image sets. Our goal is to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end,…

Cited by 186SourcePDFScholar
2015

Face Video Retrieval With Image Query via Hashing Across Euclidean Space and Riemannian Manifold

CVPR 2015poster

Retrieving videos of a specific person given his/her face image as query becomes more and more appealing for applications like smart movie fast-forwards and suspect searching. It also forms an interesting but challenging computer vision task, as the visual data to match, i.e., still image and video…

Cited by 82SourcePDFScholar
2015

Leveraging Datasets With Varying Annotations for Face Alignment via Deep Regression Network

ICCV 2015poster

Facial landmark detection, as a vital topic in computer vision, has been studied for many decades and lots of datasets have been collected for evaluation. These datasets usually have different annotations, e.g., 68-landmark markup for LFPW dataset, while 74-landmark markup for GTAV dataset. Intuitiv…

Cited by 35PDFScholar
2015

Log-Euclidean Metric Learning on Symmetric Positive Definite Manifold with Application to Image Set Classification

ICML 2015poster

The manifold of Symmetric Positive Definite (SPD) matrices has been successfully used for data representation in image set classification. By endowing the SPD manifold with Log-Euclidean Metric, existing methods typically work on vector-forms of SPD matrix logarithms. This however not only inevitabl…

Cited by 313SourcePDFScholar
2015

Projection Metric Learning on Grassmann Manifold With Application to Video Based Face Recognition

CVPR 2015poster

In video based face recognition, great success has been made by representing videos as linear subspaces, which typically lie in a special type of non-Euclidean space known as Grassmann manifold. To leverage the kernel-based methods developed for Euclidean space, several recent methods have been prop…

Cited by 297SourcePDFScholar
2015

Two Birds, One Stone: Jointly Learning Binary Code for Large-Scale Face Image Retrieval and Attributes Prediction

ICCV 2015poster

We address the challenging large-scale content-based face image retrieval problem, intended as searching images based on the presence of specific subject, given one face image of him/her. To this end, one natural demand is a supervised binary code learning method. While the learned codes might be di…

Cited by 65PDFScholar