← Search

Xiangyu Zhu

49 accepted papers

2026

AdaField: Generalizable Surface Pressure Modeling with Physics-Informed Pre-training and Flow-Conditioned Adaptation

AAAI 2026technical

The surface pressure field of transportation systems, including cars, trains, and aircraft, is critical for aerodynamic analysis and design. In recent years, deep neural networks have emerged as promising and efficient methods for modeling surface pressure field, being alternatives to computationall

Cited by 0SourcePDFScholar
2026

DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization

CVPR 2026

The rapid evolution of deepfake technologies demands robust and reliable face forgery detection algorithms. While determining whether an image has been manipulated remains essential, the ability to precisely localize forgery clues is also important for enhancing model explainability and building use

Cited by 0SourceScholar
2026

From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing

CVPR 2026

Face recognition remains vulnerable to presentation attacks, calling for robust Face Anti-Spoofing (FAS) solutions. Recent MLLM-based FAS methods reformulate the binary classification task as the generation of brief textual descriptions to improve cross-domain generalization. However, their generali

Cited by 0SourceScholar
2026

PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation

CVPR 2026

Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over talking face, such as speaking style and emotional expression, resulting in uniform facial motion. In this paper, we focus on impro

Cited by 0SourcecodeScholar
2026

Pose-RFT: Aligning MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning

ICLR 2026poster

Generating 3D human poses from multimodal inputs such as text or images requires models to capture both rich semantic and spatial correspondences. While pose-specific multimodal large language models (MLLMs) have shown promise, their supervised fine-tuning (SFT) paradigm struggles to resolve the tas…

Cited by 0SourceScholar
2026

ReGenHOI: Unifying Reconstruction and Generation for 3D Human-Object Interaction Understanding

CVPR 2026

Understanding 3D human-object interaction (HOI) involves two highly-related abilities: reconstruction, which perceives observed geometry, and generation, which imagines plausible future interactions. However, most existing methods treat these abilities as separate tasks, limiting their capacity to c

Cited by 0SourcecodeScholar
2026

STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars Reconstruction

CVPR 2026

Reconstructing high-fidelity and animatable 3D head avatars from monocular videos remains a challenging yet essential task. Existing methods based on 3D Gaussian Splatting typically bind Gaussians to mesh triangles and model deformations solely via Linear Blend Skinning, which results in rigid motio

Cited by 0SourcecodeScholar
2026

Unifying Locality of KANs and Feature Drift Compensation Projection for Data-Free Replay Based Continual Face Forgery Detection

AAAI 2026technical

The rapid advancements in face forgery techniques necessitate that detectors continuously adapt to new forgery methods, thus situating face forgery detection within a continual learning paradigm. However, when detectors learn new forgery types, their performance on previous types often degrades rapi

Cited by 0SourcePDFScholar
2025

Data Center Cooling System Optimization Using Offline Reinforcement Learning

ICLR 2025poster

The recent advances in information technology and artificial intelligence have fueled a rapid expansion of the data center (DC) industry worldwide, accompanied by an immense appetite for electricity to power the DCs. In a typical DC, around 30-40% of the energy is spent on the cooling system rather…

Cited by 0SourcePDFScholar
2025

DevFD : Developmental Face Forgery Detection by Learning Shared and Orthogonal LoRA Subspaces

NeurIPS 2025poster

The rise of realistic digital face generation and manipulation poses significant social risks. The primary challenge lies in the rapid and diverse evolution of generation techniques, which often outstrip the detection capabilities of existing models. To defend against the ever-evolving new types of…

Cited by 0SourceScholar
2025

Diffusion Models are Zero-Shot Generative Text-Vision Retrievers

ICASSP 2025accepted

Large-scale text-to-image diffusion models have demonstrated impressive capabilities for downstream tasks by leveraging strong vision-language alignment from generative pre-training. Recently, a number of works have explored how to use the power of text-to-image diffusion models for text-image match…

Cited by 0SourceScholar
2025

H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps

ICRA 2025

Solving real-world complex tasks using reinforcement learning (RL) without high-fidelity simulation environments or large amounts of offline data can be quite challenging. Online RL agents trained in imperfect simulation environments can suffer from severe sim-to-real issues. Offline RL approaches a

Cited by 17SourceScholar
2025

MVBoost: Boost 3D Reconstruction with Multi-View Refinement

CVPR 2025poster

Recent advancements in 3D object reconstruction have been remarkable, yet most current 3D models rely heavily on existing 3D datasets. The scarcity of diverse 3D datasets results in limited generalization capabilities of 3D reconstruction models. In this paper, we propose a novel framework for boost…

2025

Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Data

CVPR 2025poster

It is highly desirable to obtain a model that can generate high-quality 3D meshes from text prompts in just seconds. While recent attempts have adapted pre-trained text-to-image diffusion models, such as Stable Diffusion (SD), into generators of 3D representations (e.g., Triplane), they often suffer…

2025

Top-Down Guidance for Learning Object-Centric Representations

IJCAI 2025

Humans' innate ability to decompose scenes into objects allows for efficient understanding, predicting, and planning. In light of this, Object-Centric Learning (OCL) attempts to endow networks with similar capabilities, learning to represent scenes with the composition of objects. However, existing

2024

3D Face Reconstruction with the Geometric Guidance of Facial Part Segmentation

CVPR 2024highlight

3D Morphable Models (3DMMs) provide promising 3D face reconstructions in various applications. However existing methods struggle to reconstruct faces with extreme expressions due to deficiencies in supervisory signals such as sparse or inaccurate landmarks. Segmentation information contains effectiv…

2024

ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation

ECCV 2024poster

"By leveraging the text-to-image diffusion prior, score distillation can synthesize 3D contents without paired text-3D training data. Instead of spending hours of online optimization per text prompt, recent studies have been focused on learning a text-to-3D generative network for amortizing multiple…

2024

SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis

CVPR 2024poster

Achieving high synchronization in the synthesis of realistic speech-driven talking head videos presents a significant challenge. Traditional Generative Adversarial Networks (GAN) struggle to maintain consistent facial identity while Neural Radiance Fields (NeRF) methods although they can address thi…

2023

EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation

ICCV 2023poster

Speech-driven 3D face animation aims to generate realistic facial expressions that match the speech content and emotion. However, existing methods often neglect emotional facial expressions or fail to disentangle them from speech content. To address this issue, this paper proposes an end-to-end neur…

Cited by 117PDFcodeScholar
2023

Graphics Capsule: Learning Hierarchical 3D Face Representations From 2D Images

CVPR 2023poster

The function of constructing the hierarchy of objects is important to the visual process of the human brain. Previous studies have successfully adopted capsule networks to decompose the digits and faces into parts in an unsupervised manner to investigate the similar perception mechanism of neural ne…

Cited by 7SourcePDFScholar
2023

Grouped Knowledge Distillation for Deep Face Recognition

AAAI 2023technical

Compared with the feature-based distillation methods, logits distillation can liberalize the requirements of consistent feature dimension between teacher and student networks, while the performance is deemed inferior in face recognition. One major challenge is that the light-weight student network h…

Cited by 11SourcePDFScholar
2023

High-Fidelity Clothed Avatar Reconstruction From a Single Image

CVPR 2023poster

This paper presents a framework for efficient 3D clothed avatar reconstruction. By combining the advantages of the high accuracy of optimization-based methods and the efficiency of learning-based methods, we propose a coarse-to-fine way to realize a high-fidelity clothed avatar reconstruction (CAR)…

2023

Intrinsic Physical Concepts Discovery With Object-Centric Predictive Models

CVPR 2023poster

The ability to discover abstract physical concepts and understand how they work in the world through observing lies at the core of human intelligence. The acquisition of this ability is based on compositionally perceiving the environment in terms of objects and relations in an unsupervised manner. R…

Cited by 9SourcePDFScholar
2023

NerVE: Neural Volumetric Edges for Parametric Curve Extraction From Point Cloud

CVPR 2023poster

Extracting parametric edge curves from point clouds is a fundamental problem in 3D vision and geometry processing. Existing approaches mainly rely on keypoint detection, a challenging procedure that tends to generate noisy output, making the subsequent edge extraction error-prone. To address this is…

2023

OTAvatar: One-Shot Talking Face Avatar With Controllable Tri-Plane Rendering

CVPR 2023poster

Controllability, generalizability and efficiency are the major objectives of constructing face avatars represented by neural implicit field. However, existing methods have not managed to accommodate the three requirements simultaneously. They either focus on static portraits, restricting the represe…

2023

When Data Geometry Meets Deep Function: Generalizing Offline Reinforcement Learning

ICLR 2023poster

In offline reinforcement learning (RL), one detrimental issue to policy learning is the error accumulation of deep \textit{Q} function in out-of-distribution (OOD) areas. Unfortunately, existing offline RL methods are often over-conservative, inevitably hurting generalization performance outside dat…

2022

Constraints Penalized Q-learning for Safe Offline Reinforcement Learning

AAAI 2022technical

We study the problem of safe offline reinforcement learning (RL), the goal is to learn a policy that maximizes long-term reward while satisfying safety constraints given only offline data, without further interaction with the environment. This problem is more appealing for real world RL applications…

Cited by 105SourcePDFScholar
2022

Deconfounding Physical Dynamics with Global Causal Relation and Confounder Transmission for Counterfactual Prediction

AAAI 2022technical

Discovering the underneath causal relations is the fundamental ability for reasoning about the surrounding environment and predicting the future states in the physical world. Counterfactual prediction from visual input, which requires simulating future states based on unrealized situations in the pa…

Cited by 5SourcePDFScholar
2022

DeepThermal: Combustion Optimization for Thermal Power Generating Units Using Offline Reinforcement Learning

AAAI 2022technical

Optimizing the combustion efficiency of a thermal power generating unit (TPGU) is a highly challenging and critical task in the energy industry. We develop a new data-driven AI system, namely DeepThermal, to optimize the combustion control strategy for TPGUs. At its core, is a new model-based offlin…

Cited by 87SourcePDFScholar
2022

HP-Capsule: Unsupervised Face Part Discovery by Hierarchical Parsing Capsule Network

CVPR 2022poster

Capsule networks are designed to present the objects by a set of parts and their relationships, which provide an insight into the procedure of visual perception. Although recent works have shown the success of capsule networks on simple objects like digits, the human faces with homologous structures…

Cited by 22PDFScholar
2022

OBJECT DYNAMICS DISTILLATION FOR SCENE DECOMPOSITION AND REPRESENTATION

ICLR 2022poster

The ability to perceive scenes in terms of abstract entities is crucial for us to achieve higher-level intelligence. Recently, several methods have been proposed to learn object-centric representations of scenes with multiple objects, yet most of which focus on static scenes. In this paper, we work…

Cited by 6SourcePDFScholar
2020

Beyond 3DMM Space: Towards Fine-grained 3D Face Reconstruction

ECCV 2020poster

Recently, deep learning based 3D face reconstruction methods have shown promising results in both quality and efficiency. However, most of their training data is constructed by 3D Morphable Model, whose space spanned is only a small part of the shape space. As a result, the reconstruction results lo…

2020

Deep Spatial Gradient and Temporal Depth Learning for Face Anti-Spoofing

CVPR 2020oral

Face anti-spoofing is critical to the security of face recognition systems. Depth supervised learning has been proven as one of the most effective methods for face anti-spoofing. Despite the great success, most previous works still formulate the problem as a single-frame multi-task one by simply aug…

Cited by 245PDFcodeScholar
2020

Learning Meta Face Recognition in Unseen Domains

CVPR 2020oral

Face recognition systems are usually faced with unseen domains in real-world applications and show unsatisfactory performance due to their poor generalization. For example, a well-trained model on webface data cannot deal with the ID vs. Spot task in surveillance scenario. In this paper, we aim to l…

Cited by 189PDFcodeScholar
2020

Towards Fast, Accurate and Stable 3D Dense Face Alignment

ECCV 2020poster

Accurate and Stable 3D Dense Face Alignment","Existing methods of 3D dense face alignment mainly concentrate on accuracy, thus limiting the scope of their practical applications. In this paper, we propose a novel regression framework which makes a balance among speed, accuracy and stability. Firstly…

2019

Semantic Alignment: Finding Semantically Consistent Ground-Truth for Facial Landmark Detection

CVPR 2019poster

Recently, deep learning based facial landmark detection has achieved great success. Despite this, we notice that the semantic ambiguity greatly degrades the detection performance. Specifically, the semantic ambiguity means that some landmarks (e.g. those evenly distributed along the face contour) do…

Cited by 74PDFScholar
2019

Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian Detection

ICCV 2019poster

Multispectral pedestrian detection has shown great advantages under poor illumination conditions, since the thermal modality provides complementary information for the color image. However, real multispectral data suffers from the position shift problem, i.e. the color-thermal image pairs are not st…

Cited by 241PDFcodeScholar
2017

S3FD: Single Shot Scale-Invariant Face Detector

ICCV 2017poster

This paper presents a real-time face detector, named Single Shot Scale-invariant Face Detector (S3FD), which performs superiorly on various scales of faces with a single deep neural network, especially for small faces. Specifically, we try to solve the common problem that anchor-based detectors dete…

Cited by 887PDFcodeScholar
2015

High-Fidelity Pose and Expression Normalization for Face Recognition in the Wild

CVPR 2015poster

Pose and expression normalization is a crucial step to recover the canonical view of faces under arbitrary conditions, so as to improve the face recognition performance. An ideal normalization method is desired to be automatic, database independent and high-fidelity, where the face appearance should…

Cited by 726SourcePDFScholar
2015

Person Re-Identification by Local Maximal Occurrence Representation and Metric Learning

CVPR 2015poster

Person re-identification is an important technique towards automatic search of a person's presence in a surveillance video. Two fundamental problems are critical for person re-identification, feature representation and metric learning. An effective feature representation should be robust to illumina…

Cited by 2616SourcePDFScholar