← Search

Jiankang Deng

100 accepted papers

2026

Balancing Marker and Markerless Modes in Vision-Based Tactile Sensors with a Translucent Skin

ICRA 2026poster

Vision-based tactile sensors (VBTS) face an inherent trade-off in tactile skin design. Opaque ink markers enable accurate force and tangential displacement estimation but occlude geometric features essential for object and texture classification. Conversely, markerless skins preserve surface details…

Cited by 0Scholar
2026

CASteer: Cross-Attention Steering for Controllable Concept Erasure

ICLR 2026poster

Diffusion models have transformed image generation, yet controlling their outputs for diverse applications, including content moderation and creative customization, remains challenging. Existing approaches usually require task-specific training and struggle to generalise across both concrete (e.g.,…

Cited by 0SourcecodeScholar
2026

CEDex: Cross-Embodiment Dexterous Grasp Generation at Scale from Human-Like Contact Representations

ICRA 2026poster

Cross-embodiment dexterous grasp synthesis refers to adaptively generating and optimizing grasps for various robotic hands with different morphologies. This capability is crucial for achieving versatile robotic manipulation in diverse environments and requires substantial amounts of reliable and div…

2026

Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing

CVPR 2026

Always-on sensing is essential for next-generation edge/wearable AI systems, yet continuous high-fidelity RGB video capture remains prohibitively expensive for resource-constrained mobile and edge platforms. We present a new paradigm for efficient streaming video understanding: grayscale-always, col

Cited by 0SourcecodeScholar
2026

Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

CVPR 2026

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of gesture-rich data and their limited ability to infer fine-grained

Cited by 0SourceScholar
2026

Dynamic Novel View Synthesis in High Dynamic Range

ICLR 2026poster

High Dynamic Range Novel View Synthesis (HDR NVS) seeks to learn an HDR 3D model from Low Dynamic Range (LDR) training images captured under conventional imaging conditions. Current methods primarily focus on static scenes, implicitly assuming all scene elements remain stationary and non-living. How…

Cited by 0SourcecodeScholar
2026

FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation

CVPR 2026

The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse and precise camera trajectories. While real-world captures are photorealistic, they are typically sparse and discrete. Conversely, synthetic data sc

Cited by 0SourcecodeScholar
2026

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

ICML 2026poster

Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively exploiting these imagined futures for action execution remains challenging. Existing approaches either condition policies on predicted frames or dire…

Cited by 0SourceScholar
2026

ImagiDrive: A Unified Imagination-And-Planning Framework for Autonomous Driving

ICRA 2026poster

Autonomous driving requires rich contextual comprehension and precise predictive reasoning to navigate dynamic and complex environments safely. Vision-Language Models (VLMs) and Driving World Models (DWMs) have independently emerged as powerful recipes addressing different aspects of this challenge.…

2026

Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion Models

CVPR 2026

Generating realistic human-human interactions is a challenging task that requires not only high-quality individual body and hand motions, but also coherent coordination among all interactants. Due to limitations in available data and increased learning complexity, previous methods tend to ignore han

Cited by 0SourceScholar
2026

LiteVSR: Enabling Cross-Domain Fine-Grained Detail Generation in Light-Weight Transformers for Video Super-Resolution

ICML 2026poster

Large-scale pre-trained video generators offer powerful priors for Video Super-Resolution (VSR), yet adapting them remains computationally prohibitive. Full fine-tuning demands extensive resources, and ControlNet-style adapters lose their efficiency advantage under modern Diffusion Transformers (DiT…

Cited by 0SourceScholar
2026

MIDSTEER: Optimal Affine Framework for Steering Generative Models

ICML 2026poster

Steering intermediate representations has emerged as a powerful strategy for controlling generative models. However, despite its empirical success, it currently lacks a comprehensive theoretical framework. In this paper, we bridge this gap by formalizing the theory of concept steering. First, we est…

Cited by 0SourceScholar
2026

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

AAAI 2026technical

The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visual data, and deictic gestures, which frequently leads to task failure. Existing monolithic Vision-Language Models (VLMs)

Cited by 0SourcePDFScholar
2026

ReFocusEraser: Refocusing for Small Object Removal with Robust Context-Shadow Repair

ICLR 2026poster

Existing diffusion-based object removal and inpainting methods often fail to recover the fine structural and textural details of small objects. This is primarily due to the VAE encoder’s downsampling, which inevitably compresses small masked regions and causes significant detail loss, while the deco…

Cited by 0SourcecodeScholar
2026

SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action Planning

CVPR 2026

Vision-Language-Action (VLA) models have emerged as a promising paradigm where pretrained Vision-Language Models (VLMs) serve as System 2 for high-level reasoning, connected to action experts as System 1 for low-level motor control.However, current works fail to genuinely leverage VLM capabilities:

Cited by 0SourceScholar
2026

Towards Streaming Referring Video Segmentation via Large Language Model

CVPR 2026

Current referring video segmentation methods typically operate in an offline manner, where sparse frames are first selected for image-level referring segmentation, and the resulting masks are then propagated across the video. Although video sampling captures global context, its isolated processing s

Cited by 0SourcecodeScholar
2026

UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning

AAAI 2026technical

Universal multimodal embedding models are essential in various tasks. Existing approaches typically use in-batch mining to identify hard negatives by measuring the similarity of query-candidate pairs. However, these methods often struggle to capture subtle semantic differences among candidates and l

Cited by 0SourcePDFScholar
2026

Unleashing Vision-Language Semantics for Deepfake Video Detection

CVPR 2026

Recent Deepfake Video Detection (DFD) studies have demonstrated that pre-trained Vision-Language Models (VLMs) such as CLIP exhibit strong generalization capabilities in detecting artifacts across different identities. However, existing approaches focus on leveraging visual features only, overlookin

Cited by 0SourcecodeScholar
2026

ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs

AAAI 2026technical

Large Multimodal Models (LMMs) often face a modality representation gap during pretraining: while language embeddings remain stable, visual representations are highly sensitive to contextual noise (e.g., background clutter). To address this issue, we introduce a visual comprehension stage, which we

Cited by 0SourcePDFScholar
2025

"Principal Components" Enable A New Language of Images

ICCV 2025poster

We introduce a novel visual tokenization framework that embeds a provable PCA-like structure into the latent token space. While existing visual tokenizers primarily optimize for reconstruction fidelity, they often neglect the structural properties of the latent space--a critical factor for both inte…

2025

CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination

AAAI 2025technical

Contrastive Language-Image Pre-training (CLIP) has achieved excellent performance over a wide range of tasks. However, the effectiveness of CLIP heavily relies on a substantial corpus of pre-training data, resulting in notable consumption of computational resources. Although knowledge distillation h…

Cited by 5SourcePDFScholar
2025

CaricatureBooth: Data-Free Interactive Caricature Generation in a Photo Booth

CVPR 2025poster

We present CaricatureBooth, a system that transforms caricature creation into a simple interactive experience -- as easy as using a photo booth! A key challenge in caricature generation is two-fold: the scarcity of high-quality caricature data and the difficulty in enabling precise creative control…

2025

Deep Gaussian from Motion: Exploring 3D Geometric Foundation Models for Gaussian Splatting

NeurIPS 2025poster

Neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) are popular techniques to reconstruct and render photorealistic images. However, the prerequisite of running Structure-from-Motion (SfM) to get camera poses limits their completeness. Although previous methods can reconstruct a few unpos…

Cited by 0SourceScholar
2025

ForCenNet: Foreground-Centric Network for Document Image Rectification

ICCV 2025poster

Document image rectification aims to eliminate geometric deformation in photographed documents to facilitate text recognition. However, existing methods often neglect the significance of foreground elements, which provide essential geometric references and layout information for document image corre…

2025

Fractal Calibration for Long-tailed Object Detection

CVPR 2025poster

Real-world datasets follow an imbalanced distribution, which poses significant challenges in rare-category object detection. Recent studies tackle this problem by developing re-weighting and re-sampling methods, that utilise the class frequencies of the dataset. However, these techniques focus solel…

2025

Frequency-Guided Diffusion for Training-Free Text-Driven Image Translation

ICCV 2025poster

Current training-free text-driven image translation primarily uses diffusion features (convolution and attention) of pre-trained model as guidance to preserve the style/structure of source image in translated image. However, the coarse guidance at feature level struggles with style (e.g., visual pat…

Cited by 0SourcePDFScholar
2025

From Attention to Activation: Unraveling the Enigmas of Large Language Models

ICLR 2025poster

We study two strange phenomena in auto-regressive Transformers: (1) the dominance of the first token in attention heads; (2) the occurrence of large outlier activations in the hidden states. We find that popular large language models, such as Llama attend maximally to the first token in 98% of attenti…

2025

Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution

NeurIPS 2025poster

End-to-end autonomous driving methods aim to directly map raw sensor inputs to future driving actions such as planned trajectories, bypassing traditional modular pipelines. While these approaches have shown promise, they often operate under a one-shot paradigm that relies heavily on the current scen…

Cited by 0SourcecodeScholar
2025

HUST: High-Fidelity Unbiased Skin Tone Estimation via Texture Quantization

ICCV 2025poster

Recent 3D facial reconstruction methods have made significant progress in shape estimation, but high-fidelity unbiased facial albedo estimation remains challenging. Existing methods rely on expensive light-stage captured data, and while they have made progress in either high-fidelity reconstruction…

2025

HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos

CVPR 2025highlight

Despite the advent in 3D hand pose estimation, current methods predominantly focus on single-image 3D hand reconstruction in the camera frame, overlooking the world-space motion of the hands. Such limitation prohibits their direct use in egocentric video settings, where hands and camera are continuo…

2025

ImHead: A Large-scale Implicit Morphable Model for Localized Head Modeling

ICCV 2025poster

Over the last years, 3D morphable models (3DMMs) have emerged as a state-of-the-art methodology for modeling and generating expressive 3D avatars. However, given their reliance on a strict topology, along with their linear nature, they struggle to represent complex full-head shapes. Following the ad…

Cited by 0SourcePDFScholar
2025

Region-based Cluster Discrimination for Visual Representation Learning

ICCV 2025poster

Learning visual representations is foundational for a broad spectrum of downstream tasks. Although recent vision-language contrastive models, such as CLIP and SigLIP, have achieved impressive zero-shot performance via large-scale vision-language alignment, their reliance on global representations co…

2025

S^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion Priors

CVPR 2025poster

Recent 3D face reconstruction methods have made remarkable advancements, yet achieving high-quality facial reflectance from monocular input remains challenging. Existing methods rely on the light-stage captured data to learn facial reflectance models. However, limited subject diversity in these data…

Cited by 0SourcePDFScholar
2025

ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D Modeling

NeurIPS 2025poster

3D generation from natural language offers significant potential to reduce expert manual modeling efforts and enhance accessibility to 3D assets. However, existing methods often yield unstructured meshes and exhibit poor interactivity, making them impractical for artistic workflows. To address these…

Cited by 0SourceScholar
2025

Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator

ICCV 2025poster

Sign language is a visual language that encompasses all linguistic features of natural languages and serves as the primary communication method for the deaf and hard-of-hearing communities. Although many studies have successfully adapted pretrained language models (LMs) for sign language translation…

2025

Single-view Image to Novel-view Generation for Hand-Object Interactions

AAAI 2025technical

Hand-object interaction modeling from a single RGB image is a significantly challenging task. Previous works typically reconstruct hand-object interactions as texture-less meshes, ignoring photo-realistic image generation. In this work, we introduce the HO123, a novel method to synthesize novel-view…

Cited by 0SourcePDFScholar
2025

UniViT: Unifying Image and Video Understanding in One Vision Encoder

NeurIPS 2025poster

Despite the impressive progress of recent pretraining methods on multimodal tasks, existing methods are inherently biased towards either spatial modeling (e.g., CLIP) or temporal modeling (e.g., V-JEPA), limiting their joint capture of spatial details and temporal dynamics. To this end, we propose U…

Cited by 0SourceScholar
2025

Unlocking the Potential of Diffusion Priors in Blind Face Restoration

ICCV 2025poster

Although diffusion prior is rising as a powerful solution for blind face restoration (BFR), the inherent gap between the vanilla diffusion model and BFR settings hinders its seamless adaptation. The gap mainly stems from the discrepancy between 1) high-quality (HQ) and low-quality (LQ) images and 2)…

Cited by 0SourcePDFScholar
2025

Unsupervised Audio-Visual Segmentation with Modality Alignment

AAAI 2025technical

Audio-Visual Segmentation (AVS) aims to identify, at the pixel level, the object in a visual scene that produces a given sound. Current AVS methods rely on costly fine-grained annotations of mask-audio pairs, making them impractical for scalability. To address this, we propose the Modality Correspo…

2025

VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning

ICCV 2025poster

In recent years, video question answering based on multimodal large language models (MLLM) has garnered considerable attention, due to the benefits from the substantial advancements in LLMs. However, these models have a notable deficiency in the domains of video temporal grounding and reasoning, pos…

2025

WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild

CVPR 2025poster

In recent years, 3D hand pose estimation methods have garnered significant attention due to their extensive applications in human-computer interaction, virtual reality, and robotics. In contrast, there has been a notable gap in hand detection pipelines, posing significant challenges in constructing…

2024

3DGazeNet: Generalizing Gaze Estimation with Weak Supervision from Synthetic Views

ECCV 2024poster

"Developing gaze estimation models that generalize well to unseen domains and in-the-wild conditions remains a challenge with no known best solution. This is mostly due to the difficulty of acquiring ground truth data that cover the distribution of faces, head poses, and environments that exist in t…

Cited by 8SourcePDFScholar
2024

AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis

NeurIPS 2024poster

Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a condition for synthesizing binaural audio. However, in additio…

2024

Any-Size-Diffusion: Toward Efficient Text-Driven Synthesis for Any-Size HD Images

AAAI 2024technical

Stable diffusion, a generative model used in text-to-image synthesis, frequently encounters resolution-induced composition problems when generating images of varying sizes. This issue primarily stems from the model being trained on pairs of single-scale images and their corresponding text descriptio…

2024

Arc2Face: A Foundation Model for ID-Consistent Human Faces

ECCV 2024oral

"This paper presents , an identity-conditioned face foundation model, which, given the ArcFace embedding of a person, can generate diverse photo-realistic images with an unparalleled degree of face similarity than existing models. Despite previous attempts to decode face recognition features into de…

2024

Boosting Object Detection with Zero-Shot Day-Night Domain Adaptation

CVPR 2024poster

Detecting objects in low-light scenarios presents a persistent challenge as detectors trained on well-lit data exhibit significant performance degradation on low-light data due to low visibility. Previous methods mitigate this issue by exploring image enhancement or object detection techniques with…

2024

CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-spoofing

CVPR 2024highlight

Domain generalization (DG) based Face Anti-Spoofing (FAS) aims to improve the model's performance on unseen domains. Existing methods either rely on domain labels to align domain-invariant feature spaces or disentangle generalizable features from the whole sample which inevitably lead to the distort…

Cited by 34SourcePDFScholar
2024

DiffSED: Sound Event Detection with Denoising Diffusion

AAAI 2024technical

Sound Event Detection (SED) aims to predict the temporal boundaries of all the events of interest and their class labels, given an unconstrained audio sample. Taking either the split-and-classify (i.e., frame-level) strategy or the more principled event-level modeling approach, all existing methods…

2024

ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling

NeurIPS 2024poster

We propose ID-to-3D, a method to generate identity- and text-guided 3D human heads with disentangled expressions, starting from even a single casually captured ‘in-the-wild’ image of a subject. The foundation of our approach is anchored in compositionality, alongside the use of task-specific 2D diff…

Cited by 2SourcePDFScholar
2024

Monocular Identity-Conditioned Facial Reflectance Reconstruction

CVPR 2024poster

Recent 3D face reconstruction methods have made remarkable advancements yet there remain huge challenges in monocular high-quality facial reflectance reconstruction. Existing methods rely on a large amount of light-stage captured data to learn facial reflectance models. However the lack of subject d…

Cited by 3SourcePDFScholar
2024

Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization

NeurIPS 2024poster

The Mixture of Experts (MoE) paradigm provides a powerful way to decompose dense layers into smaller, modular computations often more amenable to human interpretation, debugging, and editability. However, a major challenge lies in the computational cost of scaling the number of experts high enough t…

2024

Neural Sign Actors: A Diffusion Model for 3D Sign Language Production from Text

CVPR 2024poster

Sign Languages (SL) serve as the primary mode of communication for the Deaf and Hard of Hearing communities. Deep learning methods for SL recognition and translation have achieved promising results. However Sign Language Production (SLP) poses a challenge as the generated motions must be realistic a…

Cited by 20SourcePDFScholar
2024

RWKV-CLIP: A Robust Vision-Language Representation Learner

EMNLP 2024main

Contrastive Language-Image Pre-training (CLIP) has significantly improved performance in various vision-language tasks by expanding the dataset with image-text pairs obtained from the web. This paper further explores CLIP from the perspectives of data and model architecture. To mitigate the impact o…

2024

SAGS: Structure-Aware 3D Gaussian Splatting

ECCV 2024poster

"Following the advent of NeRFs, 3D Gaussian Splatting (3D-GS) has paved the way to real-time neural rendering overcoming the computational burden of volumetric methods. Several extensions of 3D-GS have been proposed to achieve compressible and high-fidelity performance. However, by employing a geome…

Cited by 7SourcePDFScholar
2024

Self-Adaptive Reality-Guided Diffusion for Artifact-Free Super-Resolution

CVPR 2024poster

Artifact-free super-resolution (SR) aims to translate low-resolution images into their high-resolution counterparts with a strict integrity of the original content eliminating any distortions or synthetic details. While traditional diffusion-based SR techniques have demonstrated remarkable abilities…

2024

Three Heads Are Better than One: Complementary Experts for Long-Tailed Semi-supervised Learning

AAAI 2024technical

We address the challenging problem of Long-Tailed Semi-Supervised Learning (LTSSL) where labeled data exhibit imbalanced class distribution and unlabeled data follow an unknown distribution. Unlike in balanced SSL, the generated pseudo-labels are skewed towards head classes, intensifying the trainin…

2024

TopoFR: A Closer Look at Topology Alignment on Face Recognition

NeurIPS 2024poster

The field of face recognition (FR) has undergone significant advancements with the rise of deep learning. Recently, the success of unsupervised learning and graph neural networks has demonstrated the effectiveness of data structure information. Considering that the FR task can leverage large-scale…

2024

Unified Physical-Digital Face Attack Detection

IJCAI 2024poster

Face Recognition (FR) systems can suffer from physical (i.e., print photo) and digital (i.e., DeepFake) attacks. However, previous related work rarely considers both situations at the same time. This implies the deployment of multiple models and thus more computational burden. The main reasons for t…

Cited by 15SourcePDFScholar
2024

VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections

NeurIPS 2024poster

Large language models (LLMs) have recently emerged as powerful tools for tackling many language-processing tasks. Despite their success, training and fine-tuning these models is still far too computationally and memory intensive. In this paper, we identify and characterise the important components n…

2024

VkD: Improving Knowledge Distillation using Orthogonal Projections

CVPR 2024poster

Knowledge distillation is an effective method for training small and efficient deep learning models. However the efficacy of a single method can degenerate when transferring to other tasks modalities or even other architectures. To address this limitation we propose a novel constrained feature disti…

2023

ALIP: Adaptive Language-Image Pre-Training with Synthetic Caption

ICCV 2023poster

Contrastive Language-Image Pre-training (CLIP) has significantly boosted the performance of various vision-language tasks by scaling up the dataset with image-text pairs collected from the web. However, the presence of intrinsic noise and unmatched image-text pairs in web data can potentially affect…

Cited by 54PDFcodeScholar
2023

Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes

ICCV 2023poster

The success of deep learning models on structured data has generated significant interest in extending their application to non-Euclidean domains. In this work, we introduce a novel intrinsic operator suitable for representation learning on 3D meshes. Our operator is specifically tailored to adapt i…

Cited by 0PDFcodeScholar
2023

Controllable Person Image Synthesis with Pose-Constrained Latent Diffusion

ICCV 2023poster

Controllable person image synthesis aims at rendering a source image based on user-specified changes in body pose or appearance. Prior art approaches leverage pixel-level denoising diffusion models conditioned on the coarse skeleton via cross-attention. This leads to two limitations: low efficiency…

Cited by 25PDFcodeScholar
2023

DamoFD: Digging into Backbone Design on Face Detection

ICLR 2023poster

Face detection (FD) has achieved remarkable success over the past few years, yet, these leaps often arrive when consuming enormous computation costs. Moreover, when considering a realistic situation, i.e., building a lightweight face detector under a computation-scarce scenario, such heavy computati…

2023

DiffTAD: Temporal Action Detection with Proposal Denoising Diffusion

ICCV 2023poster

We propose a new formulation of temporal action detection (TAD) with denoising diffusion, DiffTAD in short. Taking as input random temporal proposals, it can yield action proposals accurately given an untrimmed long video. This presents a generative modeling perspective, against previous discriminat…

Cited by 43PDFcodeScholar
2023

FitMe: Deep Photorealistic 3D Morphable Model Avatars

CVPR 2023poster

In this paper, we introduce FitMe, a facial reflectance model and a differentiable rendering optimization pipeline, that can be used to acquire high-fidelity renderable human avatars from single or multiple images. The model consists of a multi-modal style-based generator, that captures facial appea…

Cited by 35SourcePDFScholar
2023

HeadSculpt: Crafting 3D Head Avatars with Text

NeurIPS 2023poster

Recently, text-guided 3D generative methods have made remarkable advancements in producing high-quality textures and geometry, capitalizing on the proliferation of large vision-language and image diffusion models. However, existing methods still struggle to create high-fidelity 3D head avatars in t…

Cited by 51SourcePDFScholar
2023

Holistic Label Correction for Noisy Multi-Label Classification

ICCV 2023poster

Multi-label classification aims to learn classification models from instances associated with multiple labels. It is pivotal to learn and utilize the label dependence among multiple labels in multi-label classification. As a result of today's big and complex data, noisy labels are inevitable, making…

Cited by 14PDFScholar
2023

Improving Fairness in Facial Albedo Estimation via Visual-Textual Cues

CVPR 2023highlight

Recent 3D face reconstruction methods have made significant advances in geometry prediction, yet further cosmetic improvements are limited by lagged albedo because inferring albedo from appearance is an ill-posed problem. Although some existing methods consider prior knowledge from illumination to i…

Cited by 6SourcePDFScholar
2023

Regularization of Polynomial Networks for Image Recognition

CVPR 2023poster

Deep Neural Networks (DNNs) have obtained impressive performance across tasks, however they still remain as black boxes, e.g., hard to theoretically analyze. At the same time, Polynomial Networks (PNs) have emerged as an alternative method with a promising performance and improved interpretability b…

2023

Spatio-temporal Prompting Network for Robust Video Feature Extraction

ICCV 2023poster

The frame quality deterioration problem is one of the main challenges in the field of video understanding. To compensate for the information loss caused by deteriorated frames, recent approaches exploit transformer-based integration modules to obtain spatio-temporal information. However, these integ…

Cited by 5PDFcodeScholar
2023

TransFace: Calibrating Transformer Training for Face Recognition from a Data-Centric Perspective

ICCV 2023poster

Vision Transformers (ViTs) have demonstrated powerful representation ability in various visual tasks thanks to their intrinsic data-hungry nature. However, we unexpectedly find that ViTs perform vulnerably when applied to face recognition (FR) scenarios with extremely large datasets. We investigate…

Cited by 34PDFcodeScholar
2023

Unicom: Universal and Compact Representation Learning for Image Retrieval

ICLR 2023poster

Modern image retrieval methods typically rely on fine-tuning pre-trained encoders to extract image-level descriptors. However, the most widely used models are pre-trained on ImageNet-1K with limited classes. The pre-trained feature representation is therefore not universal enough to generalize well…

2022

Augmenting Deep Classifiers with Polynomial Neural Networks

ECCV 2022poster

"Deep neural networks have been the driving force behind the success in classification tasks, e.g., object and audio recognition. Impressive results and generalization have been achieved by a variety of recently proposed architectures, the majority of which are seemingly disconnected. In this work,…

2022

Decoupled Multi-Task Learning With Cyclical Self-Regulation for Face Parsing

CVPR 2022poster

This paper probes intrinsic factors behind typical failure cases (e.g spatial inconsistency and boundary confusion) produced by the existing state-of-the-art method in face parsing. To tackle these problems, we propose a novel Decoupled Multi-task Learning with Cyclical Self-Regulation (DML-CSR) for…

Cited by 43PDFcodeScholar
2022

Killing Two Birds With One Stone: Efficient and Robust Training of Face Recognition CNNs by Partial FC

CVPR 2022poster

Learning discriminative deep feature embeddings by using million-scale in-the-wild datasets and margin-based softmax loss is the current state-of-the-art approach for face recognition. However, the memory and computing cost of the Fully Connected (FC) layer linearly scales up to the number of identi…

Cited by 104PDFcodeScholar
2022

Long-Tailed Instance Segmentation Using Gumbel Optimized Loss

ECCV 2022poster

"Major advancements have been made in the field of object detection and segmentation recently. However, when it comes to rare categories, the state-of-the-art methods fail to detect them, resulting in a significant performance gap between rare and frequent categories. In this paper, we identify that…

2022

MimicME: A Large Scale Diverse 4D Database for Facial Expression Analysis

ECCV 2022poster

"Recently, Deep Neural Networks (DNNs) have been shown to outperform traditional methods in many disciplines such as computer vision, speech recognition and natural language processing. A prerequisite for the successful application of DNNs is the big number of data. Even though various facial datase…

2022

MogFace: Towards a Deeper Appreciation on Face Detection

CVPR 2022poster

Benefiting from the pioneering design of generic object detectors, significant achievements have been made in the field of face detection. Typically, the architectures of the backbone, feature pyramid layer, and detection head module within the face detector all assimilate the excellent experience f…

Cited by 36PDFcodeScholar
2022

Physically-Based Face Rendering for NIR-VIS Face Recognition

NeurIPS 2022accept

Near infrared (NIR) to Visible (VIS) face matching is challenging due to the significant domain gaps as well as a lack of sufficient data for cross-modality model training. To overcome this problem, we propose a novel method for paired NIR-VIS facial image generation. Specifically, we reconstruct 3D…

2022

Sample and Computation Redistribution for Efficient Face Detection

ICLR 2022poster

Although tremendous strides have been made in uncontrolled face detection, accurate face detection with a low computation cost remains an open challenge. In this paper, we point out that computation distribution and scale augmentation are the keys to detecting small faces from low-resolution images.…

2021

Poly-NL: Linear Complexity Non-Local Layers With 3rd Order Polynomials

ICCV 2021poster

Spatial self-attention layers, in the form of Non-Local blocks, introduce long-range dependencies in Convolutional Neural Networks by computing pairwise similarities among all possible positions. Such pairwise functions underpin the effectiveness of non-local layers, but also determine a complexity…

Cited by 14PDFScholar
2021

Variational Prototype Learning for Deep Face Recognition

CVPR 2021poster

Deep face recognition has achieved remarkable improvements due to the introduction of margin-based softmax loss, in which the prototype stored in the last linear layer represents the center of each class. In these methods, training samples are enforced to be close to positive prototypes and far apar…

Cited by 100PDFScholar
2021

WebFace260M: A Benchmark Unveiling the Power of Million-Scale Deep Face Recognition

CVPR 2021poster

In this paper, we contribute a new million-scale face benchmark containing noisy 4M identities/260M faces (WebFace260M) and cleaned 2M identities/42M faces (WebFace42M) training data, as well as an elaborately designed time-constrained evaluation protocol. Firstly, we collect 4M name list and downlo…

Cited by 313PDFScholar
2020

Dual T: Reducing Estimation Error for Transition Matrix in Label-noise Learning

NeurIPS 2020poster

The transition matrix, denoting the transition relationship from clean labels to noisy labels, is essential to build statistically consistent classifiers in label-noise learning. Existing methods for estimating the transition matrix rely heavily on estimating the noisy class posterior. However, the…

Cited by 300SourcePDFScholar
2020

GP-NAS: Gaussian Process Based Neural Architecture Search

CVPR 2020poster

Neural architecture search (NAS) advances beyond the state-of-the-art in various computer vision tasks by automating the designs of deep neural networks. In this paper, we aim to address three important questions in NAS: (1) How to measure the correlation between architectures and their performances…

Cited by 67PDFScholar
2020

P-nets: Deep Polynomial Neural Networks

CVPR 2020poster

Deep Convolutional Neural Networks (DCNNs) is currently the method of choice both for generative, as well as for discriminative learning in computer vision and machine learning. The success of DCNNs can be attributed to the careful selection of their building blocks (e.g., residual blocks, rectifier…

Cited by 95PDFcodeScholar
2020

RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild

CVPR 2020poster

Though tremendous strides have been made in uncontrolled face detection, accurate and efficient 2D face alignment and 3D face reconstruction in-the-wild remain an open challenge. In this paper, we present a novel single-shot, multi-level face localisation method, named RetinaFace, which unifies face…

Cited by 1569PDFScholar
2020

Sub-center ArcFace: Boosting Face Recognition by Large-scale Noisy Web Faces

ECCV 2020poster

Margin-based deep face recognition methods (e.g. SphereFace, CosFace, and ArcFace) have achieved remarkable success in unconstrained face recognition. However, these methods are susceptible to the massive label noise in the training data and thus require laborious human effort to clean the datasets.…

2020

Synthesizing Coupled 3D Face Modalities by Trunk-Branch Generative Adversarial Networks

ECCV 2020poster

Generating realistic 3D faces is of high importance for computer graphics and computer vision applications. Generally, research on 3D face generation revolves around linear statistical models of the facial surface. Nevertheless, these models cannot represent faithfully either the facial texture or t…

2019

ArcFace: Additive Angular Margin Loss for Deep Face Recognition

CVPR 2019oral

One of the main challenges in feature learning using Deep Convolutional Neural Networks (DCNNs) for large-scale face recognition is the design of appropriate loss functions that can enhance the discriminative power. Centre loss penalises the distance between deep features and their corresponding cla…

Cited by 8474PDFcodeScholar
2019

Dense 3D Face Decoding Over 2500FPS: Joint Texture & Shape Convolutional Mesh Decoders

CVPR 2019poster

3D Morphable Models (3DMMs) are statistical models that represent facial texture and shape variations using a set of linear bases and more particular Principal Component Analysis (PCA). 3DMMs were used as statistical priors for reconstructing 3D faces from images by solving non-linear least square o…

Cited by 93PDFScholar
2018

UV-GAN: Adversarial Facial UV Map Completion for Pose-Invariant Face Recognition

CVPR 2018poster

Recently proposed robust 3D face alignment methods establish either dense or sparse correspondence between a 3D face model and a 2D facial image. The use of these methods presents new challenges as well as opportunities for facial texture analysis. In particular, by sampling the image using the fitt…