← Search

Zhen Lei

89 accepted papers

2026

CDBridge: A Cross-omics Post-training Bridge Strategy for Context-aware Biological Modeling

ICLR 2026poster

Linking genomic DNA to quantitative, context-specific expression remains a central challenge in computational biology. Current foundation models capture either tissue context or sequence features, but not both. Cross-omics systems, in turn, often overlook critical mechanisms such as alternative spli…

Cited by 0SourceScholar
2026

DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization

CVPR 2026

The rapid evolution of deepfake technologies demands robust and reliable face forgery detection algorithms. While determining whether an image has been manipulated remains essential, the ability to precisely localize forgery clues is also important for enhancing model explainability and building use

Cited by 0SourceScholar
2026

F$^2$-Assist: Multi-Phase Fetal Growth Forecast and Report Generation from Ultrasound Examination

CVPR 2026

Forecasting fetal growth from sequential ultrasound examinations is essential for personalized prenatal care. Existing medical vision-language models (MLLMs) are limited to single-phase/organ evaluations and qualitative reasoning, neglecting longitudinal history and precise continuous biometric valu

Cited by 0SourceScholar
2026

From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing

CVPR 2026

Face recognition remains vulnerable to presentation attacks, calling for robust Face Anti-Spoofing (FAS) solutions. Recent MLLM-based FAS methods reformulate the binary classification task as the generation of brief textual descriptions to improve cross-domain generalization. However, their generali

Cited by 0SourceScholar
2026

Geometric Flow Grounding: A Unified Manifold Decoupling Framework for Dynamics Discovery and Verification

ICML 2026oral

Modeling complex dynamics from observational data is fundamental to scientific discovery and artificial intelligence. However, existing approaches ranging from Neural ODEs to diffusion models are often plagued by the entanglement of static state representations and instantaneous motion, leading to a…

Cited by 0SourceScholar
2026

HDTree: Generative Modeling of Cellular Hierarchies for Robust Lineage Inference

ICML 2026poster

In single-cell research, tracing and analyzing high-throughput single-cell differentiation trajectories is crucial for understanding biological processes. Key to this is the robust modeling of hierarchical structures that govern cellular development. Traditional methods face limitations in computati…

Cited by 0SourceScholar
2026

Modeling Long-Tail Relations in the Operating Room via In-Context Multimodal Learning

ICML 2026poster

Operating room (OR) scene graph generation (SGG) enables holistic modeling of OR domains by encoding interactions among medical staff, tools, and equipment as triplet-based structured scene graphs. Although existing OR SGG methods demonstrate satisfactory overall performance, they exhibit substantia…

Cited by 0SourceScholar
2026

Multimodal Causality-Driven Representation Learning for Generalizable Medical Image Segmentation

CVPR 2026

Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their application to medical imaging remains challenging due to the high variability and complexity of medical data. Specifically, medical images often exhibit

Cited by 0SourcecodeScholar
2026

PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation

CVPR 2026

Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over talking face, such as speaking style and emotional expression, resulting in uniform facial motion. In this paper, we focus on impro

Cited by 0SourcecodeScholar
2026

Pose-RFT: Aligning MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning

ICLR 2026poster

Generating 3D human poses from multimodal inputs such as text or images requires models to capture both rich semantic and spatial correspondences. While pose-specific multimodal large language models (MLLMs) have shown promise, their supervised fine-tuning (SFT) paradigm struggles to resolve the tas…

Cited by 0SourceScholar
2026

ReGenHOI: Unifying Reconstruction and Generation for 3D Human-Object Interaction Understanding

CVPR 2026

Understanding 3D human-object interaction (HOI) involves two highly-related abilities: reconstruction, which perceives observed geometry, and generation, which imagines plausible future interactions. However, most existing methods treat these abilities as separate tasks, limiting their capacity to c

Cited by 0SourcecodeScholar
2026

STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars Reconstruction

CVPR 2026

Reconstructing high-fidelity and animatable 3D head avatars from monocular videos remains a challenging yet essential task. Existing methods based on 3D Gaussian Splatting typically bind Gaussians to mesh triangles and model deformations solely via Linear Blend Skinning, which results in rigid motio

Cited by 0SourcecodeScholar
2026

Toward Subspace-Perturbed Trajectory-Aware Backdoor Attacks in Deep Reinforcement Learning

ICML 2026poster

Deep Reinforcement Learning agents are in- creasingly used in safety-critical domains but remain vulnerable to stealthy backdoor attacks. Existing outer-loop attacks face a trade-off be- tween perceptual stealth, poisoning efficiency, and value-function consistency, often making the at- tack ineffec…

Cited by 0SourceScholar
2026

Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning

CVPR 2026

In AI-generated image detection, current cutting-edge methods typically adapt pre-trained foundation models through partial-parameter fine-tuning. However, these approaches often struggle to generalize to forgeries from unseen generators, as the fine-tuned models capture only limited patterns from t

Cited by 0SourceScholar
2026

Unifying Locality of KANs and Feature Drift Compensation Projection for Data-Free Replay Based Continual Face Forgery Detection

AAAI 2026technical

The rapid advancements in face forgery techniques necessitate that detectors continuously adapt to new forgery methods, thus situating face forgery detection within a continual learning paradigm. However, when detectors learn new forgery types, their performance on previous types often degrades rapi

Cited by 0SourcePDFScholar
2026

Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning

ICLR 2026oral

Deepfake detection remains a formidable challenge due to the evolving nature of fake content in real-world scenarios. However, existing benchmarks suffer from severe discrepancies from industrial practice, typically featuring homogeneous training sources and low-quality testing images, which hinder…

Cited by 0SourcecodeScholar
2026

VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning

ICML 2026poster

The growing capability of video generation poses escalating security risks, making reliable detection increasingly essential. In this paper, we introduce **VideoVeritas**, a framework that integrates fine-grained perception and fact-based reasoning. We observe that while current multi-modal large la…

Cited by 0SourceScholar
2025

Bayesian Test-Time Adaptation for Vision-Language Models

CVPR 2025poster

Test-time adaptation with pre-trained vision-language models, such as CLIP, aims to adapt the model to new, potentially out-of-distribution test data. Existing methods calculate the similarity between visual embedding and learnable class embeddings, which are initialized by text embeddings, for zer…

Cited by 0SourcePDFScholar
2025

DaCapo: Score Distillation as Stacked Bridge for Fast and High-quality 3D Editing

CVPR 2025poster

Score Distillation Sampling (SDS) has been successfully extended to text-driven 3D scene editing with 2D pretrained diffusion models. However, SDS-based editing methods suffer from lengthy optimization processes with slow inference and low quality. We attribute the issue of lengthy optimization to t…

Cited by 0SourcePDFScholar
2025

DevFD : Developmental Face Forgery Detection by Learning Shared and Orthogonal LoRA Subspaces

NeurIPS 2025poster

The rise of realistic digital face generation and manipulation poses significant social risks. The primary challenge lies in the rapid and diverse evolution of generation techniques, which often outstrip the detection capabilities of existing models. To defend against the ever-evolving new types of…

Cited by 0SourceScholar
2025

Diffusion Models are Zero-Shot Generative Text-Vision Retrievers

ICASSP 2025accepted

Large-scale text-to-image diffusion models have demonstrated impressive capabilities for downstream tasks by leveraging strong vision-language alignment from generative pre-training. Recently, a number of works have explored how to use the power of text-to-image diffusion models for text-image match…

Cited by 0SourceScholar
2025

FIRM: Flexible Interactive Reflection ReMoval

AAAI 2025technical

Removing reflection from a single image is challenging due to the absence of general reflection priors. Although existing methods incorporate extensive user guidance for satisfactory performance, they often lack the flexibility to adapt user guidance in different modalities, and dense user interacti…

2025

Generating Gezi Opera Scores with a Large Language Model and a High-Quality Dataset

ICASSP 2025accepted

Despite significant progress in music generation technology recently, covering various unique styles and genres, the generation of Chinese opera scores still urgently requires more attention, primarily due to the lack of a high-quality and lyric-melody alignment opera score dataset. In this study, w…

Cited by 0SourceScholar
2025

Layer-Animate for Transparent Video Generation

ICASSP 2025accepted

Transparent videos with alpha channels play a crucial role in film production, advertising, and augmented reality fields. However, there is currently no available method for producing transparent videos. Traditional methods are time-consuming and labor-intensive, and employing alternative approaches…

Cited by 0SourceScholar
2025

MVBoost: Boost 3D Reconstruction with Multi-View Refinement

CVPR 2025poster

Recent advancements in 3D object reconstruction have been remarkable, yet most current 3D models rely heavily on existing 3D datasets. The scarcity of diverse 3D datasets results in limited generalization capabilities of 3D reconstruction models. In this paper, we propose a novel framework for boost…

2025

MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization

CVPR 2025poster

Masked Image Modeling (MIM) with Vector Quantization (VQ) has achieved great success in both self-supervised pre-training and image generation. However, most existing methods struggle to address the trade-off in the shared latent space for generation quality vs. representation learning and efficienc…

2025

Mixture-of-Attack-Experts with Class Regularization for Unified Physical-Digital Face Attack Detection

AAAI 2025technical

Unified detection of digital and physical attacks in facial recognition systems has become a focal point of research in recent years. However, current multi-modal methods typically ignore the intra-class and inter-class variability across different types of attacks, leading to degraded performance.…

Cited by 0SourcePDFScholar
2025

Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Data

CVPR 2025poster

It is highly desirable to obtain a model that can generate high-quality 3D meshes from text prompts in just seconds. While recent attempts have adapted pre-trained text-to-image diffusion models, such as Stable Diffusion (SD), into generators of 3D representations (e.g., Triplane), they often suffer…

2025

RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object Detection

AAAI 2025technical

In radar-camera 3D object detection, the radar point clouds are sparse and noisy, which causes difficulties in fusing camera and radar modalities. To solve this, we introduce a novel query-based detection method named Radar-Camera Transformer (RCTrans). Specifically, we first design a Radar Dense En…

2025

RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images

AAAI 2025technical

In recent years, diffusion models have revolutionized visual generation, outperforming traditional frameworks like Generative Adversarial Networks (GANs). However, generating images of humans with realistic semantic parts, such as hands and faces, remains a significant challenge due to their intric…

2025

Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal Transport

CVPR 2025poster

Identifying multiple novel classes in an image, known as open-vocabulary multi-label recognition, is a challenging task in computer vision. Recent studies explore the transfer of powerful vision-language models such as CLIP. However, these approaches face two critical challenges: (1) The local seman…

2025

SurgPLAN++: Universal Surgical Phase Localization Network for Online and Offline Inference

ICRA 2025

Surgical phase recognition is critical for assisting surgeons in understanding surgical videos. Existing studies focused more on online surgical phase recognition, by leveraging preceding frames to predict the current frame. Despite great progress, they formulated the task as a series of frame-wise

Cited by 4SourcecodeScholar
2025

Top-Down Guidance for Learning Object-Centric Representations

IJCAI 2025

Humans' innate ability to decompose scenes into objects allows for efficient understanding, predicting, and planning. In light of this, Object-Centric Learning (OCL) attempts to endow networks with similar capabilities, learning to represent scenes with the composition of objects. However, existing

2025

Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media Manipulation

CVPR 2025poster

To tackle the threat of fake news, the task of detecting and grounding multi-modal media manipulation (DGM4) has received increasing attention. However, most state-of-the-art methods fail to explore the fine-grained consistency within local content, usually resulting in an inadequate perception of d…

2024

3D Face Reconstruction with the Geometric Guidance of Facial Part Segmentation

CVPR 2024highlight

3D Morphable Models (3DMMs) provide promising 3D face reconstructions in various applications. However existing methods struggle to reconstruct faces with extreme expressions due to deficiencies in supervisory signals such as sparse or inaccurate landmarks. Segmentation information contains effectiv…

2024

CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-spoofing

CVPR 2024highlight

Domain generalization (DG) based Face Anti-Spoofing (FAS) aims to improve the model's performance on unseen domains. Existing methods either rely on domain labels to align domain-invariant feature spaces or disentangle generalizable features from the whole sample which inevitably lead to the distort…

Cited by 34SourcePDFScholar
2024

Compositional Inversion for Stable Diffusion Models

AAAI 2024technical

Inversion methods, such as Textual Inversion, generate personalized images by incorporating concepts of interest provided by user images. However, existing methods often suffer from overfitting issues, where the dominant presence of inverted concepts leads to the absence of other desired concepts. I…

2024

Compound Text-Guided Prompt Tuning via Image-Adaptive Cues

AAAI 2024technical

Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable generalization capabilities to downstream tasks. However, existing prompt tuning based frameworks need to parallelize learnable textual inputs for all categories, suffering from massive GPU memory consumption when there is a lar…

2024

Expanding Scene Graph Boundaries: Fully Open-vocabulary Scene Graph Generation via Visual-Concept Alignment and Retention

ECCV 2024oral

"Scene Graph Generation (SGG) offers a structured representation critical in many computer vision applications. Traditional SGG approaches, however, are limited by a closed-set assumption, restricting their ability to recognize only predefined object and relation categories. To overcome this, we cat…

2024

Factorized Learning Assisted with Large Language Model for Gloss-free Sign Language Translation

COLING 2024main

Previous Sign Language Translation (SLT) methods achieve superior performance by relying on gloss annotations. However, labeling high-quality glosses is a labor-intensive task, which limits the further development of SLT. Although some approaches work towards gloss-free SLT through jointly training…

Cited by 14SourcePDFScholar
2024

General Geometry-aware Weakly Supervised 3D Object Detection

ECCV 2024poster

"3D object detection is an indispensable component for scene understanding. However, the annotation of large-scale 3D datasets requires significant human effort. To tackle this problem, many methods adopt weakly supervised 3D object detection that estimates 3D boxes by leveraging 2D boxes and scene/…

2024

Open Vocabulary 3D Scene Understanding via Geometry Guided Self-Distillation

ECCV 2024poster

"The scarcity of large-scale 3D-text paired data poses a great challenge on open vocabulary 3D scene understanding, and hence it is popular to leverage internet-scale 2D data and transfer their open vocabulary capabilities to 3D models through knowledge distillation. However, the existing distillati…

2024

ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation

ECCV 2024poster

"By leveraging the text-to-image diffusion prior, score distillation can synthesize 3D contents without paired text-3D training data. Instead of spending hours of online optimization per text prompt, recent studies have been focused on learning a text-to-3D generative network for amortizing multiple…

2024

UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models

EMNLP 2024main

Sequential decision-making refers to algorithms that take into account the dynamics of the environment, where early decisions affect subsequent decisions. With large language models (LLMs) demonstrating powerful capabilities between tasks, we can’t help but ask: Can Current LLMs Effectively Make Seq…

Cited by 2SourcePDFScholar
2024

Unified Physical-Digital Face Attack Detection

IJCAI 2024poster

Face Recognition (FR) systems can suffer from physical (i.e., print photo) and digital (i.e., DeepFake) attacks. However, previous related work rarely considers both situations at the same time. This implies the deployment of multiple models and thus more computational burden. The main reasons for t…

Cited by 15SourcePDFScholar
2024

Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object Detection

NeurIPS 2024spotlight

Serialization-based methods, which serialize the 3D voxels and group them into multiple sequences before inputting to Transformers, have demonstrated their effectiveness in 3D object detection. However, serializing 3D voxels into 1D sequences will inevitably sacrifice the voxel spatial proximity. Su…

2023

Gloss-Free Sign Language Translation: Improving from Visual-Language Pretraining

ICCV 2023poster

Sign Language Translation (SLT) is a challenging task due to its cross-domain nature, involving the translation of visual-gestural language to text. Many previous methods employ an intermediate representation,i.e., gloss sequences, to facilitate SLT, thus transforming it into a two-stage task of sig…

Cited by 62PDFcodeScholar
2023

Graphics Capsule: Learning Hierarchical 3D Face Representations From 2D Images

CVPR 2023poster

The function of constructing the hierarchy of objects is important to the visual process of the human brain. Previous studies have successfully adopted capsule networks to decompose the digits and faces into parts in an unsupervised manner to investigate the similar perception mechanism of neural ne…

Cited by 7SourcePDFScholar
2023

Grouped Knowledge Distillation for Deep Face Recognition

AAAI 2023technical

Compared with the feature-based distillation methods, logits distillation can liberalize the requirements of consistent feature dimension between teacher and student networks, while the performance is deemed inferior in face recognition. One major challenge is that the light-weight student network h…

Cited by 11SourcePDFScholar
2023

High-Fidelity Clothed Avatar Reconstruction From a Single Image

CVPR 2023poster

This paper presents a framework for efficient 3D clothed avatar reconstruction. By combining the advantages of the high accuracy of optimization-based methods and the efficiency of learning-based methods, we propose a coarse-to-fine way to realize a high-fidelity clothed avatar reconstruction (CAR)…

2023

Intrinsic Physical Concepts Discovery With Object-Centric Predictive Models

CVPR 2023poster

The ability to discover abstract physical concepts and understand how they work in the world through observing lies at the core of human intelligence. The acquisition of this ability is based on compositionally perceiving the environment in terms of objects and relations in an unsupervised manner. R…

Cited by 9SourcePDFScholar
2023

Mixture Uniform Distribution Modeling and Asymmetric Mix Distillation for Class Incremental Learning

AAAI 2023technical

Exemplar rehearsal-based methods with knowledge distillation (KD) have been widely used in class incremental learning (CIL) scenarios. However, they still suffer from performance degradation because of severely distribution discrepancy between training and test set caused by the limited storage memo…

Cited by 11SourcePDFScholar
2023

OTAvatar: One-Shot Talking Face Avatar With Controllable Tri-Plane Rendering

CVPR 2023poster

Controllability, generalizability and efficiency are the major objectives of constructing face avatars represented by neural implicit field. However, existing methods have not managed to accommodate the three requirements simultaneously. They either focus on static portraits, restricting the represe…

2023

Self-similarity Driven Scale-invariant Learning for Weakly Supervised Person Search

ICCV 2023poster

Weakly supervised person search aims to jointly detect and match persons with only bounding box annotations. Existing approaches typically focus on improving the features by exploring the relations of persons. However, scale variation problem is a more severe obstacle and under-studied that a person…

Cited by 14PDFcodeScholar
2023

Sharpness-Aware Gradient Matching for Domain Generalization

CVPR 2023poster

The goal of domain generalization (DG) is to enhance the generalization capability of the model learned from a source domain to other unseen domains. The recently developed Sharpness-Aware Minimization (SAM) method aims to achieve this goal by minimizing the sharpness measure of the loss landscape.…

2022

An Intention Prediction Based Shared Control System for Point-to-Point Navigation of a Robotic Wheelchair

RA-L 2022

Shared control approaches for robotic wheelchairs aim to provide navigation assistance to humans by utilizing robot’s intelligence in environment perception and motion planning. They can be broadly classified into two categories based on human intention prediction. Without human intention prediction

Cited by 16SourceScholar
2022

CAViT: Contextual Alignment Vision Transformer for Video Object Re-identification

ECCV 2022poster

"Video object re-identification (reID) aims at re-identifying the same object under non-overlapping cameras by matching the video tracklets with cropped video frames. The key point is how to make full use of spatio-temporal interactions to extract more accurate representation. However, there are dil…

2022

Deconfounding Physical Dynamics with Global Causal Relation and Confounder Transmission for Counterfactual Prediction

AAAI 2022technical

Discovering the underneath causal relations is the fundamental ability for reasoning about the surrounding environment and predicting the future states in the physical world. Counterfactual prediction from visual input, which requires simulating future states based on unrealized situations in the pa…

Cited by 5SourcePDFScholar
2022

Decoupling and Recoupling Spatiotemporal Representation for RGB-D-Based Motion Recognition

CVPR 2022poster

Decoupling spatiotemporal representation refers to decomposing the spatial and temporal features into dimension-independent factors. Although previous RGB-D-based motion recognition methods have achieved promising performance through the tightly coupled multi-modal spatiotemporal representation, the…

Cited by 46PDFcodeScholar
2022

HP-Capsule: Unsupervised Face Part Discovery by Hierarchical Parsing Capsule Network

CVPR 2022poster

Capsule networks are designed to present the objects by a set of parts and their relationships, which provide an insight into the procedure of visual perception. Although recent works have shown the success of capsule networks on simple objects like digits, the human faces with homologous structures…

Cited by 22PDFScholar
2022

Nested Collaborative Learning for Long-Tailed Visual Recognition

CVPR 2022poster

The networks trained on the long-tailed dataset vary remarkably, despite the same training settings, which shows the great uncertainty in long-tailed learning. To alleviate the uncertainty, we propose a Nested Collaborative Learning (NCL), which tackles the problem by collaboratively learning multip…

Cited by 121PDFcodeScholar
2022

OBJECT DYNAMICS DISTILLATION FOR SCENE DECOMPOSITION AND REPRESENTATION

ICLR 2022poster

The ability to perceive scenes in terms of abstract entities is crucial for us to achieve higher-level intelligence. Recently, several methods have been proposed to learn object-centric representations of scenes with multiple objects, yet most of which focus on static scenes. In this paper, we work…

Cited by 6SourcePDFScholar
2021

Searching for Alignment in Face Recognition

AAAI 2021technical

A standard pipeline of current face recognition frameworks consists of four individual steps: locating a face with a rough bounding box and several fiducial landmarks, aligning the face image using a pre-defined template, extracting representations and comparing. Among them, face detection, landmark…

Cited by 16SourcePDFScholar
2020

Auto-Fas: Searching Lightweight Networks for Face Anti-Spoofing

ICASSP 2020accepted

With the development of mobile devices, it is hopeful and pressing to deploy face recognition and face anti-spoofing (FAS) model on cell phone or portable devices. Most of existing face anti-spoofing methods focus on building computational costly detector for better spoofing face detection performan…

Cited by 0SourceScholar
2020

Beyond 3DMM Space: Towards Fine-grained 3D Face Reconstruction

ECCV 2020poster

Recently, deep learning based 3D face reconstruction methods have shown promising results in both quality and efficiency. However, most of their training data is constructed by 3D Morphable Model, whose space spanned is only a small part of the shape space. As a result, the reconstruction results lo…

2020

Bridging the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample Selection

CVPR 2020oral

Object detection has been dominated by anchor-based detectors for several years. Recently, anchor-free detectors have become popular due to the proposal of FPN and Focal Loss. In this paper, we first point out that the essential difference between anchor-based and anchor-free detection is actually h…

Cited by 2298PDFcodeScholar
2020

Deep Spatial Gradient and Temporal Depth Learning for Face Anti-Spoofing

CVPR 2020oral

Face anti-spoofing is critical to the security of face recognition systems. Depth supervised learning has been proven as one of the most effective methods for face anti-spoofing. Despite the great success, most previous works still formulate the problem as a single-frame multi-task one by simply aug…

Cited by 245PDFcodeScholar
2020

Exclusivity-Consistency Regularized Knowledge Distillation for Face Recognition

ECCV 2020poster

Knowledge distillation is an effective tool to compress large pre-trained Convolutional Neural Networks (CNNs) or their ensembles into models applicable to mobile and embedded devices. The success of which mainly comes from two aspects: the designed student network and the exploited knowledge. Howev…

2020

Learning Meta Face Recognition in Unseen Domains

CVPR 2020oral

Face recognition systems are usually faced with unseen domains in real-world applications and show unsatisfactory performance due to their poor generalization. For example, a well-trained model on webface data cannot deal with the ID vs. Spot task in surveillance scenario. In this paper, we aim to l…

Cited by 189PDFcodeScholar
2020

Semi-Siamese Training for Shallow Face Learning

ECCV 2020poster

Most existing public face datasets, such as MS-Celeb-1M and VGGFace2, provide abundant information in both breadth (large number of IDs) and depth (sufficient number of samples) for training. However, in many real-world scenarios of face recognition, the training dataset is limited in depth, $ extit…

2020

Towards Fast, Accurate and Stable 3D Dense Face Alignment

ECCV 2020poster

Accurate and Stable 3D Dense Face Alignment","Existing methods of 3D dense face alignment mainly concentrate on accuracy, thus limiting the scope of their practical applications. In this paper, we propose a novel regression framework which makes a balance among speed, accuracy and stability. Firstly…

2019

Semantic Alignment: Finding Semantically Consistent Ground-Truth for Facial Landmark Detection

CVPR 2019poster

Recently, deep learning based facial landmark detection has achieved great success. Despite this, we notice that the semantic ambiguity greatly degrades the detection performance. Specifically, the semantic ambiguity means that some landmarks (e.g. those evenly distributed along the face contour) do…

Cited by 74PDFScholar
2019

Unsupervised Graph Association for Person Re-Identification

ICCV 2019poster

In this paper, we propose an unsupervised graph association (UGA) framework to learn the underlying viewinvariant representations from the video pedestrian tracklets. The core points of UGA are mining the underlying cross-view associations and reducing the damage of noise associations. To this end,…

Cited by 131PDFcodeScholar
2019

Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian Detection

ICCV 2019poster

Multispectral pedestrian detection has shown great advantages under poor illumination conditions, since the thermal modality provides complementary information for the color image. However, real multispectral data suffers from the position shift problem, i.e. the color-thermal image pairs are not st…

Cited by 241PDFcodeScholar
2018

Occlusion-aware R-CNN: Detecting Pedestrians in a Crowd

ECCV 2018poster

Pedestrian detection in crowded scenes is a challenging problem since the pedestrians often gather together and occlude each other. In this paper, we propose a new occlusion-aware R-CNN (OR-CNN) to improve the detection accuracy in the crowd. Specifically, we design a new aggregation loss to enforce…

Cited by 544SourcePDFScholar
2018

Single-Shot Refinement Neural Network for Object Detection

CVPR 2018poster

For object detection, the two-stage approach (e.g., Faster R-CNN) has been achieving the highest accuracy, whereas the one-stage approach (e.g., SSD) has the advantage of high efficiency. To inherit the merits of both while overcoming their disadvantages, in this paper, we propose a novel single-sho…

2017

Exclusivity-Consistency Regularized Multi-View Subspace Clustering

CVPR 2017spotlight

Multi-view subspace clustering aims to partition a set of multi-source data into their underlying groups. To boost the performance of multi-view clustering, numerous subspace learning algorithms have been developed in recent years, but with rare exploitation of the representation complementarity bet…

Cited by 324PDFScholar
2017

S3FD: Single Shot Scale-Invariant Face Detector

ICCV 2017poster

This paper presents a real-time face detector, named Single Shot Scale-invariant Face Detector (S3FD), which performs superiorly on various scales of faces with a single deep neural network, especially for small faces. Specifically, we try to solve the common problem that anchor-based detectors dete…

Cited by 887PDFcodeScholar
2015

High-Fidelity Pose and Expression Normalization for Face Recognition in the Wild

CVPR 2015poster

Pose and expression normalization is a crucial step to recover the canonical view of faces under arbitrary conditions, so as to improve the face recognition performance. An ideal normalization method is desired to be automatic, database independent and high-fidelity, where the face appearance should…

Cited by 726SourcePDFScholar