← Search

Peipei Li

11 accepted papers

2026

AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs

CVPR 2026

The threat of Audio-Video (AV) forgery is rapidly evolving beyond human-centric deepfakes to include more diverse manipulations across complex natural scenes. However, existing benchmarks are still confined to DeepFake-based forgeries and single-granularity annotations, thus failing to capture the d

Cited by 0SourceScholar
2026

MERLIN: Building Low-SNR Robust Multimodal LLMs for Electromagnetic Signals

CVPR 2026

The paradigm of Multimodal Large Language Models (MLLMs) offers a promising blueprint for advancing the electromagnetic (EM) domain. However, prevailing approaches often deviate from the native MLLM paradigm, instead using task-specific or pipelined architectures that lead to fundamental limitations

Cited by 0SourcecodeScholar
2026

T2Agent: A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree Search

AAAI 2026technical

Real-world multimodal misinformation often arises from mixed forgery sources, requiring dynamic reasoning and adaptive verification. However, existing methods mainly rely on static pipelines and limited tool usage, limiting their ability to handle such complexity and diversity. To address this chall

Cited by 0SourcePDFScholar
2025

Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and Understanding

CVPR 2025highlight

With the rapid growth of social media and digital photography, visually appealing images have become essential for effective communication and emotional engagement. Among the factors influencing aesthetic appeal, composition--the arrangement of visual elements within a frame--plays a crucial role. I…

Cited by 0SourcePDFScholar
2025

Eye Movements as Images: A Multimodal Framework for Eye Movements Representation

ICASSP 2025accepted

Eye movements are increasingly popular for enhancing natural language processing and modeling individual states. Although specialized methods have been developed to represent eye movements for various tasks, effectively modeling the complex dynamics of eye movements and the heterogeneity with stimul…

Cited by 0SourceScholar
2024

Class-Specific Semantic Generation and Reconstruction Learning for Open Set Recognition

IJCAI 2024poster

Open set recognition is a crucial research theme for open-environment machine learning. For this problem, a common solution is to learn compact representations of known classes and identify unknown samples by measuring deviations from these known classes. However, the aforementioned methods (1) lack…

2024

QAGait: Revisit Gait Recognition from a Quality Perspective

AAAI 2024technical

Gait recognition is a promising biometric method that aims to identify pedestrians from their unique walking patterns. Silhouette modality, renowned for its easy acquisition, simple structure, sparse representation, and convenient modeling, has been widely employed in controlled in-the-lab research.…

2020

Hierarchical Face Aging through Disentangled Latent Characteristics

ECCV 2020poster

Current age datasets lie in a long-tailed distribution, which brings difficulties to describe the aging mechanism for the imbalance ages. To alleviate it, we design a novel facial age prior to guide the aging mechanism modeling. To explore the age effects on facial images, we propose a Disentangled…

Cited by 25SourcePDFScholar
2019

M2FPA: A Multi-Yaw Multi-Pitch High-Quality Dataset and Benchmark for Facial Pose Analysis

ICCV 2019poster

Facial images in surveillance or mobile scenarios often have large view-point variations in terms of pitch and yaw angles. These jointly occurred angle variations make face recognition challenging. Current public face databases mainly consider the case of yaw variations. In this paper, a new large-s…

Cited by 45PDFcodeScholar