← Search

Shuaibo Li

6 accepted papers

2026

DualScope: Capturing Critical Spatial and Temporal Cues for Distracted Driving Activity Recognition

AAAI 2026technical

Accurately recognizing distracted driving activities in real-world scenarios is essential for improving road and pedestrian safety. However, existing approaches are prone to attending to irrelevant scene context and are susceptible to interference from redundant frames, compromising their robustness

Cited by 0SourcePDFScholar
2026

Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video Detection

ICML 2026poster

Recent advances in generative video models have blurred the boundary between real and synthetic content, raising urgent concerns about digital authenticity. Multimodal large language models (MLLMs) are appealing for AI-generated video (AIGV) forensics due to their broad perceptual and reasoning capa…

Cited by 0SourceScholar
2026

ImpText: A Benchmark and Tool-Augmented Framework for Implicit Text Reasoning

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated exceptional proficiency in standard text extraction, but they encounter significant challenges when confronting real-world implicit text. Such content typically contains malicious information, intentionally concealed through physical deforma…

Cited by 0SourceScholar
2026

RealNet: Efficient and Unsupervised Detection of AI-Generated Images via Real-Only Representation Learning

AAAI 2026technical

Detecting AI-generated images remains a persistent challenge, as existing detectors often struggle to generalize to forgeries produced by previously unseen generative models. This generalization gap mainly stems from entanglement with semantic content and overfitting to model-specific artifacts. Mor

Cited by 0SourcePDFScholar
2026

SynerDetect: Hierarchical Synergistic Learning for Generalizable AI-Generated Image Detection

AAAI 2026technical

The rapid advancement of generative models, which produce increasingly realistic synthetic images, urgently demands robust and generalizable detection methods. Consequently, research has largely pivoted to leveraging large-scale Vision Foundation Models (VFMs) for enhanced generalization. However, e

Cited by 0SourcePDFScholar
2024

UnionFormer: Unified-Learning Transformer with Multi-View Representation for Image Manipulation Detection and Localization

CVPR 2024poster

We present UnionFormer a novel framework that integrates tampering clues across three views by unified learning for image manipulation detection and localization. Specifically we construct a BSFI-Net to extract tampering features from RGB and noise views achieving enhanced responsiveness to boundary…

Cited by 10SourcePDFScholar