← Search

jun lan

17 accepted papers

2026

Adaptive and Balanced Re-initialization for Long-timescale Continual Test-time Domain Adaptation

ICASSP 2026poster

Continual test-time domain adaptation (CTTA) aims to adjust models so that they can perform well over time across non-stationary environments. While previous methods have made considerable efforts to optimize the adaptation process, a crucial question remains: Can the model adapt to continually chan…

Cited by 0SourcePDFScholar
2026

FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded Reasoning

ICLR 2026poster

The rapid rise of image generation calls for detection methods that are both interpretable and reliable. Existing approaches, though accurate, act as black boxes and fail to generalize to out-of-distribution data, while multi-modal large language models (MLLMs) provide reasoning ability but often ha…

Cited by 0SourcecodeScholar
2026

Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images

CVPR 2026

The rapid growth of AI-generated imagery has blurred the boundary between real and synthetic content, raising practical concerns for digital integrity. Vision-language models (VLMs) can provide natural language explanations, but standard one-pass classifiers often miss subtle artifacts in high-quali

Cited by 0SourceScholar
2026

VTONGuard: Automatic Detection and Authentication of AI-Generated Virtual Try-On Content

ICASSP 2026poster

With the rapid advancement of generative AI, virtual try-on (VTON) systems are becoming increasingly common in e-commerce and digital entertainment. However, the growing realism of AI-generated try-on content raises pressing concerns about authenticity and responsible use. To address this, we presen…

Cited by 0SourcePDFScholar
2026

Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning

ICLR 2026oral

Deepfake detection remains a formidable challenge due to the evolving nature of fake content in real-world scenarios. However, existing benchmarks suffer from severe discrepancies from industrial practice, typically featuring homogeneous training sources and low-quality testing images, which hinder…

Cited by 0SourcecodeScholar
2026

VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning

ICML 2026poster

The growing capability of video generation poses escalating security risks, making reliable detection increasingly essential. In this paper, we introduce **VideoVeritas**, a framework that integrates fine-grained perception and fact-based reasoning. We observe that while current multi-modal large la…

Cited by 0SourceScholar
2026

Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception

ICML 2026poster

Multimodal Large Language Models (MLLMs) excel at broad visual understanding but still struggle with fine-grained perception, where decisive evidence is small and easily overwhelmed by global context. Recent "Thinking-with-Images" methods alleviate this by iteratively zooming into regions of interes…

Cited by 0SourceScholar
2025

Efficient Transfer Learning for Video-language Foundation Models

CVPR 2025poster

Pre-trained vision-language models provide a robust foundation for efficient transfer learning across various downstream tasks. In the field of video action recognition, mainstream approaches often introduce additional modules to capture temporal information. Although the additional modules increase…

2025

Stochastic Layer-Wise Shuffle for Improving Vision Mamba Training

ICML 2025poster

Recent Vision Mamba (Vim) models exhibit nearly linear complexity in sequence length, making them highly attractive for processing visual data. However, the training methodologies and their potential are still not sufficiently explored. In this paper, we investigate strategies for Vim and propose St…

2025

WildFake: A Large-Scale and Hierarchical Dataset for AI-Generated Images Detection

AAAI 2025technical

The development of text-to-image generative models has enabled the creation of images so realistic that distinguishing between AI-generated images and real photos is becoming a challenge. This progress offers new possibilities but also raises concerns over privacy, authenticity, and security. Detect…

2024

COIN-Matting: Confounder Intervention for Image Matting

ECCV 2024poster

"Deep learning methods have significantly advanced the performance of image matting. However, dataset biases can mislead the matting models to biased behavior. In this paper, we identify the two typical biases in existing matting models, specifically contrast bias and transparency bias, and discuss…

Cited by 0SourcePDFScholar
2024

ComFusion: Enhancing Personalized Generation by Instance-Scene Compositing and Fusion

ECCV 2024poster

"Recent progress in personalizing text-to-image (T2I) diffusion models has demonstrated their capability to generate images based on personalized visual concepts using only a few user-provided examples. However, these models often struggle with maintaining high visual fidelity, particularly when mod…

Cited by 1SourcePDFScholar
2024

DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric Finetuning

NeurIPS 2024poster

The recent progress in text-to-image models pretrained on large-scale datasets has enabled us to generate various images as long as we provide a text prompt describing what we want. Nevertheless, the availability of these models is still limited when we expect to generate images that fall into a spe…

2024

Segment Anything Model Meets Image Harmonization

ICASSP 2024accepted

Image harmonization is a crucial technique in image composition that aims to seamlessly match the background by adjusting the foreground of composite images. Current methods adopt either global-level or pixel-level feature matching. Global-level feature matching ignores the proximity prior, treating…

Cited by 0SourceScholar
2023

DiffUTE: Universal Text Editing Diffusion Model

NeurIPS 2023poster

Diffusion model based language-guided image editing has achieved great success recently. However, existing state-of-the-art diffusion models struggle with rendering correct text and text style during generation. To tackle this problem, we propose a universal self-supervised text editing diffusion mo…

2023

Mobile User Interface Element Detection via Adaptively Prompt Tuning

CVPR 2023poster

Recent object detection approaches rely on pretrained vision-language models for image-text alignment. However, they fail to detect the Mobile User Interface (MUI) element since it contains additional OCR information, which describes its content and function but is often ignored. In this paper, we d…

2022

XYLayoutLM: Towards Layout-Aware Multimodal Networks for Visually-Rich Document Understanding

CVPR 2022poster

Recently, various multimodal networks for Visually-Rich Document Understanding(VRDU) have been proposed, showing the promotion of transformers by integrating visual and layout information with the text embeddings. However, most existing approaches utilize the position embeddings to incorporate the s…

Cited by 105PDFScholar