← Search

Yanbo Fan

29 accepted papers

2026

MimicTalker: A Multimodal Interactive and Memory-Enhanced Framework for Real-Time Dyadic 3D Head Generation

CVPR 2026

Dyadic interactive head generation aims to synthesize realistic head motions that respond both verbally and non-verbally to an interlocutor in real-time conversation. The existing works often focus on offline scenarios, and struggle with a shallow understanding of the multimodal conversational conte

Cited by 0SourceScholar
2025

3D Gaussian Head Avatars with Expressive Dynamic Appearances by Compact Tensorial Representations

CVPR 2025poster

Recent studies have combined 3D Gaussian and 3D Morphable Models (3DMM) to construct high-quality 3D head avatars. In this line of research, existing methods either fail to capture the dynamic textures or incur significant overhead in terms of runtime speed or storage space. To this end, we propose…

Cited by 0SourcePDFScholar
2025

DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking Heads

ICCV 2025poster

In this work, we investigate the generation of high-fidelity, audio-driven 3D Gaussian talking heads from monocular videos. We present DGTalker, an innovative framework designed for real-time, high-fidelity, and 3D-aware talking head synthesis. By leveraging Gaussian generative priors and treating t…

Cited by 0SourcePDFScholar
2025

Diffusion-based Realistic Listening Head Generation via Hybrid Motion Modeling

CVPR 2025highlight

Listening head generation aims to synthesize non-verbal responsive listening head videos that naturally react to a certain speaker, for which, both realistic head movements, expressive facial expressions, and high visual qualities are expected. Previous approaches typically follow a two-stage pipeli…

Cited by 0SourcePDFScholar
2025

DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations

CVPR 2025poster

In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward…

2025

Fine-Grained 3D Gaussian Head Avatars Modeling from Static Captures via Joint Reconstruction and Registration

ICCV 2025poster

Recently, 3D head avatar modeling based on 3D Gaussians has demonstrated significant advantages in rendering quality and efficiency, given sufficient data. Some efforts have begun to train prior models on large datasets to develop generalizable 3D Gaussian head avatar modeling methods. Unfortunately…

Cited by 0SourcePDFScholar
2025

HERA: Hybrid Explicit Representation for Ultra-Realistic Head Avatars

CVPR 2025poster

We introduce a novel approach to creating ultra-realistic head avatars and rendering them in real time (\geq 30 fps at 2048 x1334 resolution). First, we propose a hybrid explicit representation that combines the advantages of two primitive based efficient rendering techniques. UV-mapped 3D mesh is u…

Cited by 0SourcePDFScholar
2023

3D GAN Inversion With Facial Symmetry Prior

CVPR 2023poster

Recently, a surge of high-quality 3D-aware GANs have been proposed, which leverage the generative power of neural rendering. It is natural to associate 3D GANs with GAN inversion methods to project a real image into the generator's latent space, allowing free-view consistent synthesis and editing, r…

Cited by 45SourcePDFScholar
2023

Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic Graphs

NeurIPS 2023poster

Most text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicitly for human motion synthesis. However, these compact text representations may overemphasize the action names at the exp…

2023

DPE: Disentanglement of Pose and Expression for General Video Portrait Editing

CVPR 2023poster

One-shot video-driven talking face generation aims at producing a synthetic talking video by transferring the facial motion from a video to an arbitrary portrait image. Head pose and facial expression are always entangled in facial motion and transferred simultaneously. However, the entanglement set…

2023

Enhancing Fine-Tuning Based Backdoor Defense with Sharpness-Aware Minimization

ICCV 2023poster

Backdoor defense, which aims to detect or mitigate the effect of malicious triggers introduced by attackers, is becoming increasingly critical for machine learning security and integrity. Fine-tuning based on benign data is a natural defense to erase the backdoor effect in a backdoored model. Howeve…

Cited by 66PDFcodeScholar
2023

High-Fidelity Facial Avatar Reconstruction From Monocular Video With Generative Priors

CVPR 2023poster

High-fidelity facial avatar reconstruction from a monocular video is a significant research problem in computer graphics and computer vision. Recently, Neural Radiance Field (NeRF) has shown impressive novel view rendering results and has been considered for facial avatar reconstruction. However, th…

2022

A Large-Scale Multiple-Objective Method for Black-Box Attack against Object Detection

ECCV 2022poster

"Recent studies have shown that detectors based on deep models are vulnerable to adversarial examples, even in the black-box scenario where the attacker cannot access the model information. Most existing attack methods aim to minimize the true positive rate, which often shows poor attack performance…

2022

Boosting Black-Box Attack With Partially Transferred Conditional Adversarial Distribution

CVPR 2022poster

This work studies black-box adversarial attacks against deep neural networks (DNNs), where the attacker can only access the query feedback returned by the attacked DNN model, while other information such as model parameters or the training datasets are unknown. One promising approach to improve atta…

Cited by 49PDFcodeScholar
2022

Boosting the Transferability of Adversarial Attacks with Reverse Adversarial Perturbation

NeurIPS 2022accept

Deep neural networks (DNNs) have been shown to be vulnerable to adversarial examples, which can produce erroneous predictions by injecting imperceptible perturbations. In this work, we study the transferability of adversarial examples, which is significant due to its threat to real-world application…

2022

High-Fidelity GAN Inversion for Image Attribute Editing

CVPR 2022poster

We present a novel high-fidelity generative adversarial network (GAN) inversion framework that enables attribute editing with image-specific details well-preserved (e.g., background, appearance, and illumination). We first analyze the challenges of high-fidelity GAN inversion from the perspective of…

Cited by 313PDFcodeScholar
2022

Stability Analysis and Generalization Bounds of Adversarial Training

NeurIPS 2022accept

In adversarial machine learning, deep neural networks can fit the adversarial examples on the training dataset but have poor generalization ability on the test set. This phenomenon is called robust overfitting, and it can be observed when adversarially training neural nets on common datasets, includ…

2022

StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGAN

ECCV 2022poster

"One-shot talking face generation aims at synthesizing a high-quality talking face video from an arbitrary portrait image, driven by a video or an audio segment. In this work, we provide a solution from a novel perspective that differs from existing frameworks. We first investigate the latent featur…

2021

DAE-GAN: Dynamic Aspect-Aware GAN for Text-to-Image Synthesis

ICCV 2021poster

Text-to-image synthesis refers to generating an image from a given text description, the key goal of which lies in photo realism and semantic consistency. Previous methods usually generate an initial image with sentence embedding and then refine it with fine-grained word embedding. Despite the signi…

Cited by 146PDFcodeScholar
2021

Parallel Rectangle Flip Attack: A Query-Based Black-Box Attack Against Object Detection

ICCV 2021poster

Object detection has been widely used in many safety-critical tasks, such as autonomous driving. However, its vulnerability to adversarial examples has not been sufficiently studied, especially under the practical scenario of black-box attacks, where the attacker can only access the query feedback o…

Cited by 80PDFScholar
2021

Random Noise Defense Against Query-Based Black-Box Attacks

NeurIPS 2021poster

The query-based black-box attacks have raised serious threats to machine learning models in many real applications. In this work, we study a lightweight defense method, dubbed Random Noise Defense (RND), which adds proper Gaussian noise to each query. We conduct the theoretical analysis about the ef…

2020

Sparse Adversarial Attack via Perturbation Factorization

ECCV 2020poster

This work studies the sparse adversarial attack, which aims to generate adversarial perturbations onto partial positions of one benign image, such that the perturbed image is incorrectly predicted by one deep neural network (DNN) model. The sparse adversarial attack involves two challenges, i.e., wh…

2019

Compressing Convolutional Neural Networks via Factorized Convolutional Filters

CVPR 2019poster

This work studies the model compression for deep convolutional neural networks (CNNs) via filter pruning. The workflow of a traditional pruning consists of three sequential stages: pre-training the original model, selecting the pre-trained filters via ranking according to a manually designed criteri…

Cited by 133PDFcodeScholar
2019

Context-Aware Feature and Label Fusion for Facial Action Unit Intensity Estimation With Partially Labeled Data

ICCV 2019poster

Facial action unit (AU) intensity estimation is a fundamental task for facial behaviour analysis. Most previous methods use a whole face image as input for intensity prediction. Considering that AUs are defined according to their corresponding local appearance, a few patch-based methods utilize imag…

Cited by 39PDFScholar
2019

Exact Adversarial Attack to Image Captioning via Structured Output Learning With Latent Variables

CVPR 2019poster

In this work, we study the robustness of a CNN+RNN based image captioning system being subjected to adversarial noises. We propose to fool an image captioning system to generate some targeted partial captions for an image polluted by adversarial noises, even the targeted captions are totally irrelev…

Cited by 64PDFcodeScholar