← Search

Chenhao Lin

21 accepted papers

2026

Beyond Pixels: Mining Compressed Domain Artifacts for Efficient AI-Generated Video Detection

ICML 2026poster

With the rapid advancement of high-fidelity video generation models, robust AI-generated video (AIGV) detection has become increasingly needed. While most AIGV detection methods operate in the decoded pixel domain, we observe that detection in the pixel domain inevitably entangles task-irrelevant se…

Cited by 0SourceScholar
2026

EchoBat: Echo-Vision Enhancement and Echo-Layered Sampling for Video LLMs Hallucination Mitigation

AAAI 2026technical

Recent advancements in multimodal large language models (MLLMs) have shown remarkable progress in video understanding. However, video MLLMs (VideoMLLMs) still suffer from hallucinations, generating nonsensical or irrelevant content. This issue partly stems from over-reliance on pre-trained knowledge

Cited by 0SourcePDFScholar
2026

Pay Less Attention to Function Words for Free Robustness of Vision-Language Models

ICLR 2026poster

To address the trade-off between robustness and performance for robust VLM, we observe that function words could incur vulnerability of VLMs against cross-modal adversarial attacks, and propose Function-word De-Attention (FDA) accordingly to mitigate the impact of function words. Similar to differen…

Cited by 0SourcecodeScholar
2026

Privacy on the Fly: A Predictive Adversarial Transformation Network for Mobile Sensor Data

AAAI 2026technical

Mobile motion sensors such as accelerometers and gyroscopes are now ubiquitously accessible by third-party apps via standard APIs. While enabling rich functionalities like activity recognition and step counting, this openness has also enabled unregulated inference of sensitive user traits, such as g

Cited by 0SourcePDFScholar
2026

When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigm

CVPR 2026

Recently, multimodal large language models (MLLMs) have emerged as a unified paradigm for language and image generation. Compared with diffusion models, MLLMs possess a much stronger capability for semantic understanding, enabling them to process more complex textual inputs and comprehend richer con

Cited by 0SourceScholar
2025

D3: Training-Free AI-Generated Video Detection Using Second-Order Features

ICCV 2025poster

The evolution of video generation techniques, such as Sora, has made it increasingly easy to produce high-fidelity AI-generated videos, raising public concern over the dissemination of synthetic content. However, existing detection methodologies remain limited by their insufficient exploration of te…

2025

Improving Adversarial Transferability on Vision Transformers via Forward Propagation Refinement

CVPR 2025poster

Vision Transformers (ViTs) have been widely applied in various computer vision and vision-language tasks. To gain insights into their robustness in practical scenarios, transferable adversarial examples on ViTs have been extensively studied. A typical approach to improving adversarial transferabilit…

2025

Improving Integrated Gradient-based Transferable Adversarial Examples by Refining the Integration Path

AAAI 2025technical

Transferable adversarial examples are known to cause threats in practical, black-box attack scenarios. A notable approach to improving transferability is using integrated gradients (IG), originally developed for model interpretability. In this paper, we find that existing IG-based attacks have limit…

2025

Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection

CVPR 2025poster

Recent studies have shown that large vision-language models (LVLMs) often suffer from the issue of object hallucinations (OH). To mitigate this issue, we introduce an efficient method that edits the model weights based on an unsafe subspace, which we call HalluSpace in this paper. With truthful and…

2025

One-Shot Face Avatar Generation in a Single Forward Pass with Identity Preservation

ICASSP 2025accepted

Face avatar generation has gained significant attention recently. With the help of the Neural Radiance Field (NeRF), existing 3D methods alleviate facial distortion in 2D methods under large pose changes. However, the state-of-the-art 3D methods still require additional optimization for generation o…

Cited by 0SourceScholar
2025

Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights

ICCV 2025poster

Developing reliable defenses against patch attacks on object detectors has attracted increasing interest. However, we identify that existing defense evaluations lack a unified and comprehensive framework, resulting in inconsistent and incomplete assessments of current methods. To address this issue,…

2025

Shining Yourself: High-Fidelity Ornaments Virtual Try-on with Diffusion Model

CVPR 2025poster

While virtual try-on for clothes and shoes with diffusion models has gained attraction, virtual try-on for ornaments, such as bracelets, rings, earrings, and necklaces, remains largely unexplored. Due to the intricate tiny patterns and repeated geometric sub-structures in most ornaments, it is much…

Cited by 1SourcePDFScholar
2025

TGDrag: Adding Semantic Control into Point-based Image Editing via Text Guidance

ICASSP 2025accepted

Controllable image generation has emerged as a cutting-edge subject of interest. Current interactive point-based image editing frameworks, such as DragGAN, achieve impressive results in fine-grained and controllable image editing. However, relying solely on point-based manipulations can lead to unin…

Cited by 0SourceScholar
2024

Breaking Semantic Artifacts for Generalized AI-generated Image Detection

NeurIPS 2024poster

With the continuous evolution of AI-generated images, the generalized detection of them has become a crucial aspect of AI security. Existing detectors have focused on cross-generator generalization, while it remains unexplored whether these detectors can generalize across different image scenes, e.…

2024

Collapse-Aware Triplet Decoupling for Adversarially Robust Image Retrieval

ICML 2024poster

Adversarial training has achieved substantial performance in defending image retrieval against adversarial examples. However, existing studies in deep metric learning (DML) still suffer from two major limitations: *weak adversary* and *model collapse*. In this paper, we address these two limitations…

2024

Physical 3D Adversarial Attacks against Monocular Depth Estimation in Autonomous Driving

CVPR 2024poster

Deep learning-based monocular depth estimation (MDE) extensively applied in autonomous driving is known to be vulnerable to adversarial attacks. Previous physical attacks against MDE models rely on 2D adversarial patches so they only affect a small localized region in the MDE map but fail under vari…

2024

Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis

IJCAI 2024poster

Detecting synthetic from real speech is increasingly crucial due to the risks of misinformation and identity impersonation. While various datasets for synthetic speech analysis have been developed, they often focus on specific areas, limiting their utility for comprehensive research. To fill this ga…

2024

TraceEvader: Making DeepFakes More Untraceable via Evading the Forgery Model Attribution

AAAI 2024technical

In recent few years, DeepFakes are posing serve threats and concerns to both individuals and celebrities, as realistic DeepFakes facilitate the spread of disinformation. Model attribution techniques aim at attributing the adopted forgery models of DeepFakes for provenance purposes and providing expl…

Cited by 7SourcePDFScholar
2023

Learning Heuristically-Selected and Neurally-Guided Feature for Age Group Recognition Using Unconstrained Smartphone Interaction

IJCAI 2023poster

Owing to the boom of smartphone industries, the expansion of phone users has also been significant. Besides adults, children and elders have also begun to join the population of daily smartphone users. Such an expansion indeed facilitates the further exploration of the versatility and flexibility of…

Cited by 1SourcePDFScholar
2020

BlueMemo: Depression Analysis through Twitter Posts

IJCAI 2020poster

The use of social media runs through our lives, and users' emotions are also affected by it. Previous studies have reported social organizations and psychologists using social media to find depressed patients. However, due to the variety of content published by users, it isn't effortless for the sys…

Cited by 0SourcePDFScholar
2020

RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition

ECCV 2020poster

The attention-based encoder-decoder framework has recently achieved impressive results for scene text recognition, and many variants have emerged with improvements in recognition quality. However, it performs poorly on contextless texts (e.g., random character sequences) which is unacceptable in mos…