← Search

Xiao Guo

8 accepted papers

2026

FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition

CVPR 2026

Model fusion is a key strategy for robust recognition in unconstrained scenarios, as different models provide complementary strengths. This is especially important for whole-body human recognition, where biometric cues such as face, gait, and body shape vary across samples and are typically integrat

Cited by 7SourcecodeScholar
2025

Point Cloud Self-supervised Learning via 3D to Multi-view Masked Learner

ICCV 2025poster

Recently, multi-modal masked autoencoders (MAE) has been introduced in 3D self-supervised learning, offering enhanced feature learning by leveraging both 2D and 3D data to capture richer cross-modal representations. However, these approaches have two limitations: (1) they inefficiently require both…

Cited by 0SourcePDFScholar
2025

Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector

CVPR 2025poster

Deepfake detection is a long-established research topic vital for mitigating the spread of malicious misinformation. Unlike prior methods that provide either binary classification results or textual explanations separately, we introduce a novel method capable of generating both simultaneously. Our m…

2024

Common Sense Reasoning for Deep Fake Detection

ECCV 2024poster

"State-of-the-art deepfake detection approaches rely on image-based features extracted via neural networks. While these approaches trained in a supervised manner extract likely fake features, they may fall short in representing unnatural ‘non-physical’ semantic facial attributes – blurry hairlines,…

2024

On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection

NeurIPS 2024poster

Large numbers of synthesized videos from diffusion models pose threats to information security and authenticity, leading to an increasing demand for generated content detection. However, existing video-level detection algorithms primarily focus on detecting facial forgeries and often fail to identif…

2023

Bidirectional Optical Flow NeRF: High Accuracy and High Quality under Fewer Views

AAAI 2023technical

Neural Radiance Fields (NeRF) can implicitly represent 3D-consistent RGB images and geometric by optimizing an underlying continuous volumetric scene function using a sparse set of input views, which has greatly benefited view synthesis tasks. However, NeRF fails to estimate correct geometry when gi…

Cited by 7SourcePDFScholar
2023

Hierarchical Fine-Grained Image Forgery Detection and Localization

CVPR 2023poster

Differences in forgery attributes of images generated in CNN-synthesized and image-editing domains are large, and such differences make a unified image forgery detection and localization (IFDL) challenging. To this end, we present a hierarchical fine-grained formulation for IFDL representation learn…

2022

Multi-Domain Learning for Updating Face Anti-Spoofing Models

ECCV 2022poster

"In this work, we study multi-domain learning for face anti-spoofing (MD-FAS), where a pre-trained FAS model needs to be updated to perform equally well on both source and target domains while only using target domain data for updating. We present a new model for MD-FAS, which addresses the forgetti…