← Search

Yunfei Liu

22 accepted papers

2025

AnyTalk: Multi-modal Driven Multi-domain Talking Head Generation

AAAI 2025technical

Cross-domain talking head generation, such as animating a static cartoon animal photo with real human video, is crucial for personalized content creation. However, prior works typically rely on domain-specific frameworks and paired videos, limiting its utility and complicating its architecture with…

Cited by 0SourcePDFScholar
2025

CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space Modulation

ICCV 2025poster

Video face swapping aims to address two primary challenges: effectively transferring the source identity to the target video and accurately preserving the dynamic attributes of the target face, such as head poses, facial expressions, lip-sync, etc..Existing methods mainly focus on achieving high-qua…

Cited by 0SourcePDFScholar
2025

GUAVA: Generalizable Upper Body 3D Gaussian Avatar

ICCV 2025poster

Reconstructing a high-quality, animatable 3D human avatar with expressive facial and hand motions from a single image has gained significant attention due to its broad application potential. 3D human avatar reconstruction typically requires multi-view or monocular videos and training on individual I…

Cited by 0SourcePDFScholar
2025

HRAvatar: High-Quality and Relightable Gaussian Head Avatar

CVPR 2025poster

Reconstructing animatable and high-quality 3D head avatars from monocular videos, especially with realistic relighting, is a valuable task. However, the limited information from single-view input, combined with the complex head poses and facial movements, makes this challenging. Previous methods ach…

Cited by 0SourcePDFScholar
2025

SKRAG: A Retrieval-Augmented Generation Framework Guided by Reasoning Skeletons over Knowledge Graphs

EMNLP 2025

In specialized domains such as space science and utilization, question answering (QA) systems are required to perform complex multi-fact reasoning over sparse knowledge graphs (KGs). Existing KG-based retrieval-augmented generation (RAG) frameworks often face challenges such as inefficient subgraph

Cited by 0SourcePDFScholar
2025

TEASER: Token Enhanced Spatial Modeling for Expressions Reconstruction

ICLR 2025poster

3D facial reconstruction from a single in-the-wild image is a crucial task in human-centered computer vision tasks. While existing methods can recover accurate facial shapes, there remains significant space for improvement in fine-grained expression capture. Current approaches struggle with irregul…

Cited by 1SourcePDFScholar
2024

A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing

CVPR 2024poster

This paper presents a video inversion approach for zero-shot video editing which models the input video with low-rank representation during the inversion process. The existing video editing methods usually apply the typical 2D DDIM inversion or naive spatial-temporal DDIM inversion before editing wh…

2024

AddMe: Zero-shot Group-photo Synthesis by Inserting People into Scenes

ECCV 2024poster

"While large text-to-image diffusion models have made significant progress in high-quality image generation, challenges persist when users insert their portraits into existing photos, especially group photos. Concretely, existing customization methods struggle to insert facial identities at desired…

2024

DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation

CVPR 2024poster

We propose DiffSHEG a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation. While previous works focused on co-speech gesture or expression generation individually the joint generation of synchronized expressions and gestures remains barely explored. To address th…

Cited by 37SourcePDFScholar
2024

GPAvatar: Generalizable and Precise Head Avatar from Image(s)

ICLR 2024poster

Head avatar reconstruction, crucial for applications in virtual reality, online meetings, gaming, and film industries, has garnered substantial attention within the computer vision community. The fundamental objective of this field is to faithfully recreate the head avatar and precisely control expr…

2024

Revealing Hierarchical Structure of Leaf Venations in Plant Science via Label-Efficient Segmentation: Dataset and Method

IJCAI 2024poster

Hierarchical leaf vein segmentation is a crucial but under-explored task in agricultural sciences, where analysis of the hierarchical structure of plant leaf venation can contribute to plant breeding. While current segmentation techniques rely on data-driven models, there is no publicly available da…

2023

Accurate 3D Face Reconstruction with Facial Component Tokens

ICCV 2023poster

Accurately reconstructing 3D faces from monocular images and videos is crucial for various applications, such as digital avatar creation. However, the current deep learning-based methods face significant challenges in achieving accurate reconstruction with disentangled facial parameters and ensuring…

Cited by 23PDFScholar
2023

MODA: Mapping-Once Audio-driven Portrait Animation with Dual Attentions

ICCV 2023poster

Audio-driven portrait animation aims to synthesize portrait videos that are conditioned by given audio. Animating high-fidelity and multimodal video portraits has a variety of applications. Previous methods have attempted to capture different motion modes and generate high-fidelity portrait videos b…

Cited by 27PDFcodeScholar
2021

Generalizing Gaze Estimation With Outlier-Guided Collaborative Adaptation

ICCV 2021poster

Deep neural networks have significantly improved appearance-based gaze estimation accuracy. However, it still suffers from unsatisfactory performance when generalizing the trained model to new domains, e.g., unseen environments or persons. In this paper, we propose a plug-and-play gaze adaptation fr…

Cited by 71PDFcodeScholar
2020

Aggregating Crowd Wisdom with Side Information via a Clustering-based Label-aware Autoencoder

IJCAI 2020poster

Aggregating crowd wisdom infers true labels for objects, from multiple noisy labels provided by various sources. Besides labels from sources, side information such as object features is also introduced to achieve higher inference accuracy. Usually, the learning-from-crowds framework is adopted. Howe…

2020

Improving Knowledge Tracing via Pre-training Question Embeddings

IJCAI 2020poster

Knowledge tracing (KT) defines the task of predicting whether students can correctly answer questions based on their historical response. Although much research has been devoted to exploiting the question information, plentiful advanced information among questions and skills hasn't been well extract…

2020

Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks

ECCV 2020poster

Recent studies have shown that DNNs can be compromised by backdoor attacks crafted at training time. A backdoor attack installs a backdoor into the victim model by injecting a backdoor pattern into a small set of training examples. The victim model behaves normally on clean test data, yet consistent…

Cited by 719SourcePDFScholar
2020

Unsupervised Learning for Intrinsic Image Decomposition From a Single Image

CVPR 2020poster

Intrinsic image decomposition, which is an essential task in computer vision, aims to infer the reflectance and shading of the scene. It is challenging since it needs to separate one image into two components. To tackle this, conventional methods introduce various priors to constrain the solution, y…

Cited by 133PDFcodeScholar