← Search

Zhifeng Xie

13 accepted papers

2026

FilmSceneDesigner: Chaining Set Design for Procedural Film Scene Generation

AAAI 2026technical

Film set design plays a pivotal role in cinematic storytelling and shaping the visual atmosphere. However, the traditional process depends on expert-driven manual modeling, which is labor-intensive and time-consuming. To address this issue, we introduce FilmSceneDesigner, an automated scene generati

Cited by 0SourcePDFScholar
2026

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips

CVPR 2026

Foley art plays a pivotal role in enhancing immersive auditory experiences in film, yet manual creation of spatio-temporal aligned audio remains labor-intensive. We propose FoleyDesigner, a novel framework inspired by professional Foley workflows, integrating film clip analysis, spatio-temporal cont

Cited by 0SourceScholar
2026

GardenDesigner: Encoding Aesthetic Principles into Jiangnan Garden Construction via a Chain of Agents

CVPR 2026

Jiangnan gardens, a prominent style of Chinese classical gardens, hold great potential as digital assets for film and game production and digital tourism. However, manual modeling of Jiangnan gardens heavily relies on expert experience for layout design and asset creation, making the process time-co

Cited by 0SourcecodeScholar
2025

FilmComposer: LLM-Driven Music Production for Silent Film Clips

CVPR 2025poster

In this work, we implement music production for silent film clips using LLM-driven method. Given the strong professional demands of film music production, we propose the FilmComposer, simulating the actual workflows of professional musicians. FilmComposer is the first to combine large generative mod…

2025

HieraFashDiff: Hierarchical Fashion Design with Multi-stage Diffusion Models

AAAI 2025technical

Fashion design is a challenging and complex process. Recent works on fashion generation and editing are all agnostic of the actual fashion design process, which limits their usage in practice. In this paper, we propose a novel hierarchical diffusion-based framework tailored for fashion design, coine…

2025

LMTalker: Sparse Landmark-guided Gaussian Splatting for High-fidelity Talking Head Synthesis

ICASSP 2025accepted

3D Gaussian splatting (3DGS) has demonstrated significant potential in audio-driven talking head synthesis. However, despite notable advancements in speed and fidelity, current methods still face challenges such as inaccurate lip movements and facial artifacts. To address these issues, we propose LM…

Cited by 0SourceScholar
2023

High-Fidelity Generalized Emotional Talking Face Generation With Multi-Modal Emotion Space Learning

CVPR 2023poster

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing to handle unseen emotion styles due to limited semantics. T…

Cited by 46SourcePDFScholar
2022

ColorFormer: Image Colorization via Color Memory Assisted Hybrid-Attention Transformer

ECCV 2022poster

"Automatic image colorization is a challenging task that attracts a lot of research interest. Previous methods employing deep neural networks have produced impressive results. However, these colorization images are still unsatisfactory and far from practical applications. The reason is that semantic…

Cited by 64SourcePDFScholar
2022

HifiHead: One-Shot High Fidelity Neural Head Synthesis with 3D Control

IJCAI 2022poster

We propose HifiHead, a high fidelity neural talking head synthesis method, which can well preserve the source image's appearance and control the motion (e.g., pose, expression, gaze) flexibly with 3D morphable face models (3DMMs) parameters derived from a driving image or indicated by users. Existin…

2022

Learning To Restore 3D Face From In-the-Wild Degraded Images

CVPR 2022poster

In-the-wild 3D face modelling is a challenging problem as the predicted facial geometry and texture suffer from a lack of reliable clues or priors, when the input images are degraded. To address such a problem, in this paper we propose a novel Learning to Restore (L2R) 3D face framework for unsuperv…

Cited by 3PDFScholar
2022

Physically-Guided Disentangled Implicit Rendering for 3D Face Modeling

CVPR 2022poster

This paper presents a novel Physically-guided Disentangled Implicit Rendering (PhyDIR) framework for high-fidelity 3D face modeling. The motivation comes from two observations: widely-used graphics renderers yield excessive approximations against photo-realistic imaging, while neural rendering metho…

Cited by 8PDFScholar
2018

Recurrent Neural Networks for Automatic Replay Spoofing Attack Detection

ICASSP 2018accepted

In order to enhance the security of automatic speaker verification (ASV) systems, automatic spoofing attack detection, which discriminates the fake audio recordings from genuine human speech, has gain much attention recently. Among various ways of spoofing attacks, replay attacks are one of the most…

Cited by 0SourceScholar