← Search

Guoying Zhao

29 accepted papers

2026

EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models

ICLR 2026poster

Emotion understanding is a critical yet challenging task. Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced their capabilities in this area. However, MLLMs often suffer from ``hallucinations'', generating irrelevant or nonsensical content. To the best of our kn…

Cited by 0SourcecodeScholar
2025

Deep Change Monitoring: A Hyperbolic Representative Learning Framework and a Dataset for Long-term Fine-grained Tree Change Detection

CVPR 2025highlight

In environmental protection, tree monitoring plays an essential role in maintaining and improving ecosystem health. However, precise monitoring is challenging because existing datasets fail to capture continuous fine-grained changes in trees due to low-resolution images and high acquisition costs. I…

2025

Enhancing Facial Privacy Protection via Weakening Diffusion Purification

CVPR 2025poster

The rapid growth of social media has led to the widespread sharing of individual portrait images, which pose serious privacy risks due to the capabilities of automatic face recognition (AFR) systems for mass surveillance. Hence, protecting facial privacy against unauthorized AFR systems is essential…

2025

FreeNet: Liberating Depth-Wise Separable Operations for Building Faster Mobile Vision Architectures

AAAI 2025technical

In the pursuit of efficient vision architectures, substantial efforts have been devoted to optimizing operator efficiency. Depth-wise separable operators, such as DWConv, are found cheap in both FLOPs and parameters. As a result, they are increasingly incorporated into efficient backbones, trading f…

Cited by 0SourcePDFScholar
2025

From Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-Identification

CVPR 2025poster

Aiming to match pedestrian images captured under varying lighting conditions, visible-infrared person re-identification (VI-ReID) has drawn intensive research attention and achieved promising results. However, in real-world surveillance contexts, data is distributed across multiple devices/entities,…

2025

Learning Binary-Antithetical Information Bottleneck for Generalizable Face Anti-Spoofing

ICASSP 2025accepted

We investigate generalizable face anti-spoofing (FAS) using information bottleneck theory. As generalizable FAS aims to detect spoofing in unseen scenarios, it has recently gained significant attention. Existing methods often use adversarial strategies or auxiliary modules to learn domain-invariant…

Cited by 0SourceScholar
2024

Differentiable Auxiliary Learning for Sketch Re-Identification

AAAI 2024technical

Sketch re-identification (Re-ID) seeks to match pedestrians' photos from surveillance videos with corresponding sketches. However, we observe that existing works still have two critical limitations: (i) cross- and intra-modality discrepancies hinder the extraction of modality-shared features, (ii) s…

Cited by 10SourcePDFScholar
2024

Domain Shifting: A Generalized Solution for Heterogeneous Cross-Modality Person Re-Identification

ECCV 2024poster

"Cross-modality person re-identification (ReID) is a challenging task that aims to match cross-modality pedestrian images across multiple camera views. Existing methods are tailored to specific tasks and perform well for visible-infrared or visible-sketch ReID. However, the performance exhibits a no…

Cited by 6SourcePDFScholar
2024

PFStorer: Personalized Face Restoration and Super-Resolution

CVPR 2024poster

Recent developments in face restoration have achieved remarkable results in producing high-quality and lifelike outputs. The stunning results however often fail to be faithful with respect to the identity of the person as the models lack necessary context. In this paper we explore the potential of p…

Cited by 6SourcePDFScholar
2023

LART: Neural Correspondence Learning with Latent Regularization Transformer for 3D Motion Transfer

NeurIPS 2023poster

3D motion transfer aims at transferring the motion from a dynamic input sequence to a static 3D object and outputs an identical motion of the target with high-fidelity and realistic visual effects. In this work, we propose a novel 3D Transformer framework called LART for 3D motion transfer. With car…

2023

Modality Unifying Network for Visible-Infrared Person Re-Identification

ICCV 2023poster

Visible-infrared person re-identification (VI-ReID) is a challenging task due to large cross-modality discrepancies and intra-class variations. Existing methods mainly focus on learning modality-shared representations by embedding different modalities into the same feature space. As a result, the le…

Cited by 58PDFScholar
2022

Geometry-Contrastive Transformer for Generalized 3D Pose Transfer

AAAI 2022technical

We present a customized 3D mesh Transformer model for the pose transfer task. As the 3D pose transfer essentially is a deformation procedure dependent on the given meshes, the intuition of this work is to perceive the geometric inconsistency between the given meshes with the powerful self-attention…

2022

Learning Optimal K-Space Acquisition and Reconstruction Using Physics-Informed Neural Networks

CVPR 2022poster

The inherent slow imaging speed of Magnetic Resonance Image (MRI) has spurred the development of various acceleration methods, typically through heuristically undersampling of the associated measurement domain known as k-space. Recently, deep neural networks have been applied to reconstruct undersam…

Cited by 28PDFScholar
2022

Looking Back on Learned Experiences For Class/task Incremental Learning

ICLR 2022spotlight

Classical deep neural networks are limited in their ability to learn from emerging streams of training data. When trained sequentially on new or evolving tasks, their performance degrades sharply, making them inappropriate in real-world use cases. Existing methods tackle it by either storing old dat…

Cited by 51SourcePDFScholar
2022

PhysFormer: Facial Video-Based Physiological Measurement With Temporal Difference Transformer

CVPR 2022poster

Remote photoplethysmography (rPPG), which aims at measuring heart activities and physiological signals from facial video without any contact, has great potential in many applications. Recent deep learning approaches focus on mining subtle rPPG clues using convolutional neural networks with limited s…

Cited by 248PDFcodeScholar
2021

Dual-Cross Central Difference Network for Face Anti-Spoofing

IJCAI 2021poster

Face anti-spoofing (FAS) plays a vital role in securing face recognition systems. Recently, central difference convolution (CDC) has shown its excellent representation capacity for the FAS task via leveraging local gradient features. However, aggregating central difference clues from all neighbors/d…

2021

Intrinsic-Extrinsic Preserved GANs for Unsupervised 3D Pose Transfer

ICCV 2021poster

With the strength of deep generative models, 3D pose transfer regains intensive research interests in recent years. Existing methods mainly rely on a variety of constraints to achieve the pose transfer over 3D meshes, e.g., the need for manually encoding for shape and pose disentanglement. In this p…

Cited by 33PDFcodeScholar
2021

Non-contact Pain Recognition from Video Sequences with Remote Physiological Measurements Prediction

IJCAI 2021poster

Automatic pain recognition is paramount for medical diagnosis and treatment. The existing works fall into three categories: assessing facial appearance changes, exploiting physiological cues, or fusing them in a multi-modal manner. However, (1) appearance changes are easily affected by subjective fa…

Cited by 11SourcePDFScholar
2021

iMiGUE: An Identity-Free Video Dataset for Micro-Gesture Understanding and Emotion Analysis

CVPR 2021poster

We introduce a new dataset for the emotional artificial intelligence research: identity-free video dataset for micro-gesture understanding and emotion analysis (iMiGUE). Different from existing public datasets, iMiGUE focuses on nonverbal body gestures without using any identity information, while t…

Cited by 117PDFcodeScholar
2020

Auto-Fas: Searching Lightweight Networks for Face Anti-Spoofing

ICASSP 2020accepted

With the development of mobile devices, it is hopeful and pressing to deploy face recognition and face anti-spoofing (FAS) model on cell phone or portable devices. Most of existing face anti-spoofing methods focus on building computational costly detector for better spoofing face detection performan…

Cited by 0SourceScholar
2020

Face Anti-Spoofing with Human Material Perception

ECCV 2020poster

Face anti-spoofing (FAS) plays a vital role in securing the face recognition systems from presentation attacks. Most existing FAS methods capture various cues (e.g., texture, depth and reflection) to distinguish the live faces from the spoofing faces. All these cues are based on the discrepancy amon…

Cited by 193SourcePDFScholar
2020

Searching Central Difference Convolutional Networks for Face Anti-Spoofing

CVPR 2020poster

Face anti-spoofing (FAS) plays a vital role in face recognition systems. Most state-of-the-art FAS methods 1) rely on stacked convolutions and expert-designed network, which is weak in describing detailed fine-grained information and easily being ineffective when the environment varies (e.g., differ…

Cited by 620PDFcodeScholar
2020

Video-based Remote Physiological Measurement via Cross-verified Feature Disentangling

ECCV 2020poster

Remote physiological measurements, e.g., remote photoplethysmography (rPPG) based heart rate (HR), heart rate variability (HRV) and respiration frequency (RF) measuring, are playing more and more important roles under the application scenarios where contact measurement is inconvenient or impossible.…

2019

Remote Heart Rate Measurement From Highly Compressed Facial Videos: An End-to-End Deep Learning Solution With Video Enhancement

ICCV 2019poster

Remote photoplethysmography (rPPG), which aims at measuring heart activities without any contact, has great potential in many applications (e.g., remote healthcare). Existing rPPG approaches rely on analyzing very fine details of facial videos, which are prone to be affected by video compression. He…

Cited by 370PDFcodeScholar
2019

Structured Modeling of Joint Deep Feature and Prediction Refinement for Salient Object Detection

ICCV 2019poster

Recent saliency models extensively explore to incorporate multi-scale contextual information from Convolutional Neural Networks (CNNs). Besides direct fusion strategies, many approaches introduce message-passing to enhance CNN features or predictions. However, the messages are mainly transmitted in…

Cited by 59PDFScholar
2018

Super Wide Regression Network for Unsupervised Cross-Database Facial Expression Recognition

ICASSP 2018accepted

Unsupervised cross-database facial expression recognition (FER) is a challenging problem, in which the training and testing samples belong to different facial expression databases. For this reason, the training (source) and testing (target) facial expression samples would have different feature dist…

Cited by 0SourceScholar
2018

Unsupervised Cross-Corpus Speech Emotion Recognition Using Domain-Adaptive Subspace Learning

ICASSP 2018accepted

In this paper, we investigate an interesting problem, i.e., unsupervised cross-corpus speech emotion recognition (SER), in which the training and testing speech signals come from two different speech emotion corpora. Meanwhile, the training speech signals are labeled, while the label information of…

Cited by 0SourceScholar
2017

SRN: Side-output Residual Network for Object Symmetry Detection in the Wild

CVPR 2017oral

In this paper, we establish a baseline for object symmetry detection in complex backgrounds by presenting a new benchmark and an end-to-end deep learning approach, opening up a promising direction for symmetry detection in the wild. The new benchmark, named Sym-PASCAL, spans challenges including obj…

Cited by 120PDFcodeScholar