← Search

Zhenan Sun

33 accepted papers

2026

Artificial Immune System of Secure Face Recognition Against Adversarial Attacks (Abstract Reprint)

AAAI 2026technical

Deep learning-based face recognition models are vulnerable to adversarial attacks. In contrast to general noises, the presence of imperceptible adversarial noises can lead to catastrophic errors in deep face recognition models. The primary difference between adversarial noise and general noise lies

Cited by 0SourcePDFScholar
2026

MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts

ICML 2026poster

LLMs hold great promise for healthcare applications, but fast-changing medical knowledge can quickly make their outputs outdated or inaccurate, limiting use in high-stakes settings. Model editing can update LLMs without full retraining, but parameter-based methods often break locality and are risky …

Cited by 0SourceScholar
2026

Modeling Long-Tail Relations in the Operating Room via In-Context Multimodal Learning

ICML 2026poster

Operating room (OR) scene graph generation (SGG) enables holistic modeling of OR domains by encoding interactions among medical staff, tools, and equipment as triplet-based structured scene graphs. Although existing OR SGG methods demonstrate satisfactory overall performance, they exhibit substantia…

Cited by 0SourceScholar
2026

TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models

CVPR 2026

Vision-Language Models (VLMs), such as CLIP, have achieved impressive zero-shot recognition performance but remain highly susceptible to adversarial perturbations, posing significant risks in safety-critical scenarios. Previous training-time defenses rely on adversarial fine-tuning, which requires l

Cited by 0SourcecodeScholar
2026

UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception

AAAI 2026technical

The remarkable success of diffusion models in text-to-image generation has sparked growing interest in expanding their capabilities to a variety of multi-modal tasks, including image understanding, manipulation, and perception. These tasks require advanced semantic comprehension across both visual a

Cited by 0SourcePDFScholar
2026

VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction

AAAI 2026technical

Generating responsive listener head dynamics with nuanced emotions and expressive reactions is crucial for dialogue modeling in various virtual avatar animations. Previous studies mainly focus on the direct short-term production of listener behavior. They overlook the fine-grained control over motio

Cited by 0SourcePDFScholar
2025

Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance

ICASSP 2025accepted

Text-editable and pose-controllable character video generation is a challenging but prevailing topic with practical applications. However, existing approaches mainly focus on single-object video generation with pose guidance, ignoring the realistic situation that multi-character appear concurrently…

Cited by 0SourceScholar
2025

Revealing Key Details to See Differences: A Novel Prototypical Perspective for Skeleton-based Action Recognition

CVPR 2025highlight

In skeleton-based action recognition, a key challenge is distinguishing between actions with similar trajectories of joints due to the lack of image-level details in skeletal representations. Recognizing that the differentiation of similar actions relies on subtle motion details in specific body par…

2024

DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs

NeurIPS 2024oral

Quantization of large language models (LLMs) faces significant challenges, particularly due to the presence of outlier activations that impede efficient low-bit representation. Traditional approaches predominantly address Normal Outliers, which are activations across all tokens with relatively large…

2024

Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular Images

AAAI 2024technical

Reconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held objects. Though recent works have employed implicit functions to achieve impressive…

2024

MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric

CVPR 2024poster

Vision-language pre-trained models have achieved impressive performance on various downstream tasks. However their large model sizes hinder their utilization on platforms with limited computational resources. We find that directly using smaller pre-trained models and applying magnitude-based pruning…

Cited by 21SourcePDFScholar
2022

Disentangled Federated Learning for Tackling Attributes Skew via Invariant Aggregation and Diversity Transferring

ICML 2022spotlight

Attributes skew hinders the current federated learning (FL) frameworks from consistent optimization directions among the clients, which inevitably leads to performance reduction and unstable convergence. The core problems lie in that: 1) Domain-specific attributes, which are non-causal and only loca…

2022

Mimic Embedding via Adaptive Aggregation: Learning Generalizable Person Re-identification

ECCV 2022poster

"Domain generalizable (DG) person re-identification (ReID) aims to test across unseen domains without access to the target domain data at training time, which is a realistic but challenging problem. In contrast to methods assuming an identical model for different domains, Mixture of Experts (MoE) ex…

2021

PyMAF: 3D Human Pose and Shape Regression With Pyramidal Mesh Alignment Feedback Loop

ICCV 2021poster

Regression-based methods have recently shown promising results in reconstructing human meshes from monocular images. By directly mapping raw pixels to model parameters, these methods can produce parametric models in a feed-forward manner via neural networks. However, minor deviation in parameters ma…

Cited by 386PDFcodeScholar
2021

ReMix: Towards Image-to-Image Translation With Limited Data

CVPR 2021poster

Image-to-image (I2I) translation methods based on generative adversarial networks (GANs) typically suffer from overfitting when limited training data is available. In this work, we propose a data augmentation method (ReMix) to tackle this issue. We interpolate training samples at the feature level a…

Cited by 37PDFcodeScholar
2020

A Lightweight Multi-Label Segmentation Network for Mobile Iris Biometrics

ICASSP 2020accepted

This paper proposes a novel, lightweight deep convolutional neural network specifically designed for iris segmentation of noisy images acquired by mobile devices. Unlike previous studies, which only focused on improving the accuracy of segmentation mask using the popular CNN technology, our method i…

Cited by 0SourceScholar
2020

Hierarchical Face Aging through Disentangled Latent Characteristics

ECCV 2020poster

Current age datasets lie in a long-tailed distribution, which brings difficulties to describe the aging mechanism for the imbalance ages. To alleviate it, we design a novel facial age prior to guide the aging mechanism modeling. To explore the age effects on facial images, we propose a Disentangled…

Cited by 25SourcePDFScholar
2020

Informative Sample Mining Network for Multi-Domain Image-to-Image Translation

ECCV 2020poster

The performance of multi-domain image-to-image translation has been significantly improved by recent progress in deep generative models. Existing approaches can use a unified model to achieve translations between all the visual domains. However, their outcomes are far from satisfying when there are…

Cited by 9SourcePDFScholar
2019

Distant Supervised Centroid Shift: A Simple and Efficient Approach to Visual Domain Adaptation

CVPR 2019poster

Conventional domain adaptation methods usually resort to deep neural networks or subspace learning to find invariant representations across domains. However, most deep learning methods highly rely on large-size source domains and are computationally expensive to train, while subspace learning method…

Cited by 130PDFScholar
2019

Foreground-Aware Pyramid Reconstruction for Alignment-Free Occluded Person Re-Identification

ICCV 2019poster

Re-identifying a person across multiple disjoint camera views is important for intelligent video surveillance, smart retailing and many other applications. However, existing person re-identification methods are challenged by the ubiquitous occlusion over persons and suffer performance degradation. T…

Cited by 252PDFScholar
2019

M2FPA: A Multi-Yaw Multi-Pitch High-Quality Dataset and Benchmark for Facial Pose Analysis

ICCV 2019poster

Facial images in surveillance or mobile scenarios often have large view-point variations in terms of pitch and yaw angles. These jointly occurred angle variations make face recognition challenging. Current public face databases mainly consider the case of yaw variations. In this paper, a new large-s…

Cited by 45PDFcodeScholar
2018

Deep Spatial Feature Reconstruction for Partial Person Re-Identification: Alignment-Free Approach

CVPR 2018poster

Partial person re-identification (re-id) is a challenging problem, where only a partial observation of a person image is available for matching. However, few studies have offered a solution of how to identify an arbitrary patch of a person image. In this paper, we propose a fast and accurate matchin…

2018

End-to-end View Synthesis for Light Field Imaging with Pseudo 4DCNN

ECCV 2018poster

Limited angular resolution has become the main bottleneck of microlens-based plenoptic cameras towards practical vision applications. Existing view synthesis methods mainly break the task into two steps, i.e. depth estimating and view warping, which are usually inefficient and produce artifacts over…

Cited by 139SourcePDFScholar
2018

IntroVAE: Introspective Variational Autoencoders for Photographic Image Synthesis

NeurIPS 2018poster

We present a novel introspective variational autoencoder (IntroVAE) model for synthesizing high-resolution photographic images. IntroVAE is capable of self-evaluating the quality of its generated samples and improving itself accordingly. Its inference and generator models are jointly trained in an i…

Cited by 356SourcePDFScholar
2018

Learning a High Fidelity Pose Invariant Model for High-resolution Face Frontalization

NeurIPS 2018poster

Face frontalization refers to the process of synthesizing the frontal view of a face from a given profile. Due to self-occlusion and appearance distortion in the wild, it is extremely challenging to recover faithful results and preserve texture details in a high-resolution. This paper proposes a Hi…

Cited by 113SourcePDFScholar