← Search

Zitong YU

41 accepted papers

2026

Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection

CVPR 2026

Face forgery detection faces a critical challenge: a persistent gap between offline benchmarks and real-world efficacy, which we attribute to the ecological invalidity of training data. This work introduces Agent4FaceForgery to address two fundamental problems: (1) how to capture the diverse intents

Cited by 0SourceScholar
2026

FLOW: Optimal Transport-Driven Feature Warping for Generalized Remote Physiological Measurement

CVPR 2026

Remote photoplethysmography (rPPG) enables non-contact physiological measurement from facial videos but often suffers from severe performance degradation under domain shifts. Traditional STMap-based methods [??] rely on predefined spatio-temporal representations that offer engineered robustness but

Cited by 0SourceScholar
2026

FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models

AAAI 2026technical

Face anti-spoofing (FAS) is crucial for protecting facial recognition systems from presentation attacks. Previous methods approached this task as a classification problem, lacking interpretability and reasoning behind the predicted results. Recently, multimodal large language models (MLLMs) have sho

Cited by 0SourcePDFScholar
2026

H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation

AAAI 2026technical

Unified video and action prediction models hold great potential for robotic manipulation, as future observations offer contextual cues for planning, while actions reveal how interactions shape the environment. However, most existing approaches treat observation and action generation in a monolithic

Cited by 0SourcePDFScholar
2026

Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning

ICLR 2026poster

Vision token pruning has proven to be an effective acceleration technique for the Efficient Vision Language Model (VLM). However, existing pruning methods demonstrate excellent performance preservation in visual question answering (VQA) and suffer substantial degradation on visual grounding (VG) tas…

Cited by 0SourcecodeScholar
2026

PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning

AAAI 2026technical

In recent years, face anti-spoofing (FAS) has made notable progress in multimodal fusion, cross-domain generalization, and interpretability. With the development of large language models and reinforcement learning (RL), strategy-based training paradigms offer new opportunities for jointly modeling m

Cited by 0SourcePDFScholar
2026

PHASE-Net: Physics-Grounded Harmonic Attention System for Efficient Remote Photoplethysmography Measurement

CVPR 2026

Remote photoplethysmography (rPPG) measurement enables non-contact physiological monitoring but suffers from accuracy degradation under head motion and illumination changes. Existing deep learning methods are mostly heuristic and lack theoretical grounding, limiting robustness and interpretability.

Cited by 0SourcecodeScholar
2026

PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing

ICLR 2026poster

Remote photoplethysmography (rPPG) enables non-contact physiological measurement but remains highly susceptible to illumination changes, motion artifacts, and limited temporal modeling. Large Language Models (LLMs) excel at capturing long-range dependencies, offering a potential solution but struggl…

Cited by 0SourceScholar
2026

SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition

AAAI 2026technical

Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However, when treating LLM as a recognizer, two questions arise: 1) How can LLMs underst

Cited by 0SourcePDFScholar
2026

Unsupervised Camouflaged Object Detection with Dual-Eigenvector Spectral Pseudo-Labeling and Contrastive Refinement

ICML 2026poster

Unsupervised Camouflaged Object Detection (UCOD) aims to identify objects concealed in their surroundings without relying on pixel-level labels. Existing methods rely solely on simple post-processing of DINO high-dimensional features to generate pseudo labels for training. However, these methods suf…

Cited by 0SourceScholar
2026

When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?

AAAI 2026technical

Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, which simulates an “Audio-Visual Confusion” scene by modifying the corresponding sound of an object in the video, e.g., mute

Cited by 0SourcePDFScholar
2025

Big-Moe: Bypassing Isolated Gating For Generalized Multimodal Face Anti-Spoofing

ICASSP 2025accepted

In the domain of facial recognition security, multimodal Face Anti-Spoofing (FAS) is essential for countering presentation attacks. However, existing technologies encounter challenges due to modality biases and imbalances, as well as domain shifts. Our research introduces a Mixture of Experts (MoE)…

Cited by 0SourceScholar
2025

CA-Edit: Causality-Aware Condition Adapter for High-Fidelity Local Facial Attribute Editing

AAAI 2025technical

For efficient and high-fidelity local facial attribute editing, most existing editing methods either require additional fine-tuning for different editing effects or tend to affect beyond the editing regions. Alternatively, inpainting methods can edit the target image region while preserving external…

2025

DADM: Dual Alignment of Domain and Modality for Face Anti-spoofing

ICCV 2025poster

With the availability of diverse sensor modalities (i.e., RGB, Depth, Infrared) and the success of multi-modal learning, multi-modal face anti-spoofing (FAS) has emerged as a prominent research focus. The intuition behind it is that leveraging multiple modalities can uncover more intrinsic spoofing…

2025

Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units

EMNLP 2025

This paper investigates the enhancement of reasoning capabilities in language models through token-level multi-model collaboration. Our approach selects the optimal tokens from the next token distributions provided by multiple models to perform autoregressive reasoning. Contrary to the assumption th

2025

EPE-P: Evidence-based Parameter-efficient Prompting for Multimodal Learning with Missing Modalities

ICASSP 2025accepted

Missing modalities are a common challenge in real-world multimodal learning scenarios, occurring during both training and testing. Existing methods for managing missing modalities often require the design of separate prompts for each modality or missing case, leading to complex designs and a substan…

Cited by 0SourceScholar
2025

FSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understanding

CVPR 2025poster

Figure skating, known as the "Art on Ice," is among the most artistic sports, challenging to understand due to its blend of technical elements (like jumps and spins) and overall artistic expression. Existing figure skating datasets mainly focus on single tasks, such as action recognition or scoring,…

2025

Few-Shot Audio-Visual Class-Incremental Learning with Temporal Prompting and Regularization

AAAI 2025technical

Audio-Visual Learning (AVL) aims at the audio-visual perception with both audio and vision modalities. AVL also suffers from data insufficiency in many applications as with other unimodal tasks. Concurrently, AVL often needs to continuously learn over time rather than all knowledge simultaneously. C…

Cited by 0SourcePDFScholar
2025

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners

ICLR 2025poster

Contrastive language-image pretraining (CLIP) has significantly advanced image-based vision learning. A pressing topic subsequently arises: how can we effectively adapt CLIP to the video domain? Recent studies have focused on adjusting either the textual or visual branch of CLIP for action recogniti…

2025

MSAmba: Exploring Multimodal Sentiment Analysis with State Space Models

AAAI 2025technical

Multimodal sentiment analysis, which learns a model to process multiple modalities simultaneously and predict a sentiment value, is an important area of affective computing. Modeling sequential intra-modal information and enhancing cross-modal interactions are crucial to multimodal sentiment analysi…

2025

MoEdit: On Learning Quantity Perception for Multi-object Image Editing

CVPR 2025poster

Multi-object images are widely present in the real world, spanning various areas of daily life. Efficient and accurate editing of these images is crucial for applications such as augmented reality, advertisement design, and medical imaging. Stable Diffusion (SD) has ushered in a new era of high-qual…

2025

PGD-Imp: Rethinking and Unleashing Potential of Classic PGD with Dual Strategies for Imperceptible Adversarial Attacks

ICASSP 2025accepted

Imperceptible adversarial attacks have recently attracted increasing research interests. Existing methods typically incorporate external modules or loss terms other than a simple l<inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">p</inf>-norm into the att…

Cited by 0SourceScholar
2024

MTaDCS: Moving Trace and Feature Density-based Confidence Sample Selection under Label Noise

ECCV 2024poster

"Learning from noisy labels is a challenging task, as noisy labels can compromise decision boundaries and result in suboptimal generalization performance. Most previous approaches for dealing noisy labels are based on sample selection, which utilized the small loss criterion to reduce the adverse ef…

2023

Audio-Visual Deception Detection: DOLOS Dataset and Parameter-Efficient Crossmodal Learning

ICCV 2023poster

Deception detection in conversations is a challenging yet important task, having pivotal applications in many fields such as credibility assessment in business, multimedia anti-frauds, and custom security. Despite this, deception detection research is hindered by the lack of high-quality deception d…

Cited by 12PDFcodeScholar
2023

Learning Motion-Robust Remote Photoplethysmography through Arbitrary Resolution Videos

AAAI 2023technical

Remote photoplethysmography (rPPG) enables non-contact heart rate (HR) estimation from facial videos which gives significant convenience compared with traditional contact-based measurements. In the real-world long-term health monitoring scenario, the distance of the participants and their head movem…

2023

Rehearsal-Free Domain Continual Face Anti-Spoofing: Generalize More and Forget Less

ICCV 2023oral

Face Anti-Spoofing (FAS) is recently studied under the continual learning setting, where the FAS models are expected to evolve after encountering data from new domains. However, existing methods need extra replay buffers to store previous data for rehearsal, which becomes infeasible when previous da…

Cited by 25PDFcodeScholar
2022

Domain Generalization via Shuffled Style Assembly for Face Anti-Spoofing

CVPR 2022poster

With diverse presentation attacks emerging continually, generalizable face anti-spoofing (FAS) has drawn growing attention. Most existing methods implement domain generalization (DG) on the complete representations. However, different image statistics may have unique properties for the FAS tasks. In…

Cited by 203PDFcodeScholar
2022

Geometry-Contrastive Transformer for Generalized 3D Pose Transfer

AAAI 2022technical

We present a customized 3D mesh Transformer model for the pose transfer task. As the 3D pose transfer essentially is a deformation procedure dependent on the given meshes, the intuition of this work is to perceive the geometric inconsistency between the given meshes with the powerful self-attention…

2022

IDPT: Interconnected Dual Pyramid Transformer for Face Super-Resolution

IJCAI 2022poster

Face Super-resolution (FSR) task works for generating high-resolution (HR) face images from the corresponding low-resolution (LR) inputs, which has received a lot of attentions because of the wide application prospects. However, due to the diversity of facial texture and the difficulty of reconstruc…

Cited by 19SourcePDFScholar
2022

PhysFormer: Facial Video-Based Physiological Measurement With Temporal Difference Transformer

CVPR 2022poster

Remote photoplethysmography (rPPG), which aims at measuring heart activities and physiological signals from facial video without any contact, has great potential in many applications. Recent deep learning approaches focus on mining subtle rPPG clues using convolutional neural networks with limited s…

Cited by 248PDFcodeScholar
2021

Dual-Cross Central Difference Network for Face Anti-Spoofing

IJCAI 2021poster

Face anti-spoofing (FAS) plays a vital role in securing face recognition systems. Recently, central difference convolution (CDC) has shown its excellent representation capacity for the FAS task via leveraging local gradient features. However, aggregating central difference clues from all neighbors/d…

2021

Non-contact Pain Recognition from Video Sequences with Remote Physiological Measurements Prediction

IJCAI 2021poster

Automatic pain recognition is paramount for medical diagnosis and treatment. The existing works fall into three categories: assessing facial appearance changes, exploiting physiological cues, or fusing them in a multi-modal manner. However, (1) appearance changes are easily affected by subjective fa…

Cited by 11SourcePDFScholar
2021

Pixel Difference Networks for Efficient Edge Detection

ICCV 2021poster

Recently, deep Convolutional Neural Networks (CNNs) can achieve human-level performance in edge detection with the rich and abstract edge representation capacities. However, the high performance of CNN based edge detection is achieved with a large pretrained CNN backbone, which is memory and energy…

Cited by 452PDFcodeScholar
2021

iMiGUE: An Identity-Free Video Dataset for Micro-Gesture Understanding and Emotion Analysis

CVPR 2021poster

We introduce a new dataset for the emotional artificial intelligence research: identity-free video dataset for micro-gesture understanding and emotion analysis (iMiGUE). Different from existing public datasets, iMiGUE focuses on nonverbal body gestures without using any identity information, while t…

Cited by 117PDFcodeScholar
2020

Auto-Fas: Searching Lightweight Networks for Face Anti-Spoofing

ICASSP 2020accepted

With the development of mobile devices, it is hopeful and pressing to deploy face recognition and face anti-spoofing (FAS) model on cell phone or portable devices. Most of existing face anti-spoofing methods focus on building computational costly detector for better spoofing face detection performan…

Cited by 0SourceScholar
2020

Deep Spatial Gradient and Temporal Depth Learning for Face Anti-Spoofing

CVPR 2020oral

Face anti-spoofing is critical to the security of face recognition systems. Depth supervised learning has been proven as one of the most effective methods for face anti-spoofing. Despite the great success, most previous works still formulate the problem as a single-frame multi-task one by simply aug…

Cited by 245PDFcodeScholar
2020

Face Anti-Spoofing with Human Material Perception

ECCV 2020poster

Face anti-spoofing (FAS) plays a vital role in securing the face recognition systems from presentation attacks. Most existing FAS methods capture various cues (e.g., texture, depth and reflection) to distinguish the live faces from the spoofing faces. All these cues are based on the discrepancy amon…

Cited by 193SourcePDFScholar
2020

Searching Central Difference Convolutional Networks for Face Anti-Spoofing

CVPR 2020poster

Face anti-spoofing (FAS) plays a vital role in face recognition systems. Most state-of-the-art FAS methods 1) rely on stacked convolutions and expert-designed network, which is weak in describing detailed fine-grained information and easily being ineffective when the environment varies (e.g., differ…

Cited by 620PDFcodeScholar
2020

Video-based Remote Physiological Measurement via Cross-verified Feature Disentangling

ECCV 2020poster

Remote physiological measurements, e.g., remote photoplethysmography (rPPG) based heart rate (HR), heart rate variability (HRV) and respiration frequency (RF) measuring, are playing more and more important roles under the application scenarios where contact measurement is inconvenient or impossible.…

2019

Remote Heart Rate Measurement From Highly Compressed Facial Videos: An End-to-End Deep Learning Solution With Video Enhancement

ICCV 2019poster

Remote photoplethysmography (rPPG), which aims at measuring heart activities without any contact, has great potential in many applications (e.g., remote healthcare). Existing rPPG approaches rely on analyzing very fine details of facial videos, which are prone to be affected by video compression. He…

Cited by 370PDFcodeScholar