← Search

Xuannan Liu

13 accepted papers

2026

AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs

CVPR 2026

The threat of Audio-Video (AV) forgery is rapidly evolving beyond human-centric deepfakes to include more diverse manipulations across complex natural scenes. However, existing benchmarks are still confined to DeepFake-based forgeries and single-granularity annotations, thus failing to capture the d

Cited by 0SourceScholar
2026

MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents

ICLR 2026poster

The Model Context Protocol (MCP) standardizes how large language model (LLM) agents discover, describe, and call external tools. While MCP unlocks broad interoperability, it also enlarges the attack surface by making tools first-class, composable objects with natural-language metadata, and standardi…

Cited by 0SourcecodeScholar
2026

T2Agent: A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree Search

AAAI 2026technical

Real-world multimodal misinformation often arises from mixed forgery sources, requiring dynamic reasoning and adaptive verification. However, existing methods mainly rely on static pipelines and limited tool usage, limiting their ability to handle such complexity and diversity. To address this chall

Cited by 0SourcePDFScholar
2025

Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and Understanding

CVPR 2025highlight

With the rapid growth of social media and digital photography, visually appealing images have become essential for effective communication and emotional engagement. Among the factors influencing aesthetic appeal, composition--the arrangement of visual elements within a frame--plays a crucial role. I…

Cited by 0SourcePDFScholar
2025

MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs

ICLR 2025poster

Current multimodal misinformation detection (MMD) methods often assume a single source and type of forgery for each sample, which is insufficient for real-world scenarios where multiple forgery sources coexist. The lack of a benchmark for mixed-source misinformation has hindered progress in this fie…

Cited by 0SourcePDFScholar
2025

Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs

NeurIPS 2025poster

The increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily focus on model vulnerabilities exposed by static image inputs, ignoring the temporal dynamics of video that may induce…

Cited by 0SourcecodeScholar
2024

Enhancing Generalization Of Invisible Facial Privacy Cloak Via Gradient Accumulation

ICASSP 2024accepted

The blooming of social media and face recognition (FR) systems has increased people’s concern about privacy and security. A new type of adversarial privacy cloak (class-universal) can be applied to all the images of regular users, to prevent malicious FR systems from acquiring their identity informa…

Cited by 0SourceScholar
2024

Faceptor: A Generalist Model for Face Perception

ECCV 2024oral

"With the comprehensive research conducted on various face analysis tasks, there is a growing interest among researchers to develop a unified approach to face perception. Existing methods mainly discuss unified representation and training, which lack task extensibility and application efficiency. To…

2024

InstaStyle: Inversion Noise of a Stylized Image is Secretly a Style Adviser

ECCV 2024poster

"Stylized text-to-image generation focuses on creating images from textual descriptions while adhering to a style specified by reference images. However, subtle style variations within different reference images can hinder the model from accurately learning the target style. In this paper, we propos…

2024

Localize, Understand, Collaborate: Semantic-Aware Dragging via Intention Reasoner

NeurIPS 2024poster

Flexible and accurate drag-based editing is a challenging task that has recently garnered significant attention. Current methods typically model this problem as automatically learning "how to drag" through point dragging and often produce one deterministic estimation, which presents two key limitati…

2024

Open-Set Facial Expression Recognition

AAAI 2024technical

Facial expression recognition (FER) models are typically trained on datasets with a fixed number of seven basic classes. However, recent research works (Cowen et al. 2021; Bryant et al. 2022; Kollias 2023) point out that there are far more expressions than the basic ones. Thus, when these models are…

Cited by 4SourcePDFScholar
2023

Enhancing Generalization of Universal Adversarial Perturbation through Gradient Aggregation

ICCV 2023poster

Deep neural networks are vulnerable to universal adversarial perturbation (UAP), an instance-agnostic perturbation capable of fooling the target model for most samples. Compared to instance-specific adversarial examples, UAP is more challenging as it needs to generalize across various samples and mo…

Cited by 30PDFcodeScholar
2023

Leave No Stone Unturned: Mine Extra Knowledge for Imbalanced Facial Expression Recognition

NeurIPS 2023poster

Facial expression data is characterized by a significant imbalance, with most collected data showing happy or neutral expressions and fewer instances of fear or disgust. This imbalance poses challenges to facial expression recognition (FER) models, hindering their ability to fully understand various…

Cited by 24SourcePDFScholar