← Search

Jiantao Zhou

46 accepted papers

2026

AEGIS: Adversarial Target–Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models

ICLR 2026poster

Concept erasure helps stop diffusion models (DMs) from generating harmful content; but current methods face robustness-retention trade-off. **Robustness** means the model fine-tuned by concept erasure methods resists reactivation of erased concepts, even under semantically related prompts. **Retenti…

Cited by 0SourcecodeScholar
2026

Deferred Poisoning: Making the Model More Vulnerable via Hessian Singularization

AAAI 2026technical

Recent studies have shown that deep learning models are very vulnerable to poisoning attacks. Many defense methods have been proposed to address this issue. However, traditional poisoning attacks are not as threatening as commonly believed. This is because they often cause differences in how the mod

Cited by 0SourcePDFScholar
2026

Editprint: General Digital Image Forensics via Editing Fingerprint with Self-Augmentation Training

CVPR 2026

Digital image forensics can ensure information credibility in tasks like camera source identification (CSI), synthetic image detection (SID), and social network provenance (SNP). These tasks typically rely on image processing history clues left by in-camera operations, post-capture editing, or synth

Cited by 0SourcecodeScholar
2026

Enhanced Latent-Space Adversarial Training for Super-Resolution

ICML 2026poster

Real-world super-resolution (SR) is challenging due to complex degradations. HYPIR, a recent state-of-the-art diffusion-based restoration model, struggles to deal with this task in a single step. Although a naive two-step cascade improves the results, over-saturation, limited fine-grained details, a…

Cited by 0SourceScholar
2026

Forensic-Friendly Image Manipulation via Controllable Latent Diffusion

CVPR 2026

With diffusion models demonstrating superior capabilities in image editing, more users now rely on online servers for content manipulation via textual prompts rather than traditional offline tools. Despite servers attempting to prevent the proliferation of maliciously edited content via active defen

Cited by 0SourcecodeScholar
2026

LSP Framework: A Compensatory Model for Defeating Trigger Reverse Engineering via Label Smoothing Poisoning

ICASSP 2026oral

Deep neural networks are vulnerable to backdoor attacks. Among the existing backdoor defense methods, trigger reverse engineering based approaches, which reconstruct the backdoor triggers via optimizations, are the most versatile and effective ones compared to other types of methods. In this paper,…

Cited by 0SourcePDFScholar
2026

Phy-CoSF: Physics-Guided Continuous Spectral Fields Reconstruction and Spectral Super-Resolution for Snapshot Compressive Imaging

ICML 2026poster

Recent advances have demonstrated that coded aperture snapshot spectral imaging (CASSI) systems show great potential for capturing 3D hyperspectral images (HSIs) from a single 2D measurement. Despite the inherent spectral continuity of scenes captured by CASSI, most existing reconstruction methods a…

Cited by 0SourceScholar
2026

SimpleGVR: A Simple Baseline for Latent-Cascaded Generative Video Super-Resolution

ICLR 2026poster

Cascaded pipelines, which use a base text-to-video (T2V) model for low-resolution content and a video super-resolution (VSR) model for high-resolution details, are a prevailing strategy for efficient video synthesis. However, current works suffer from two key limitations: an inefficient pixel-space…

Cited by 0SourcecodeScholar
2026

Universal Adversarial Purification with DDIM Metric Loss for Stable Diffusion

AAAI 2026technical

Stable Diffusion (SD) often produces degraded outputs when the training dataset contains adversarial noise. Adversarial purification offers a promising solution by removing adversarial noise from contaminated data. However, existing purification methods are primarily designed for classification task

Cited by 0SourcePDFScholar
2026

Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning Framework

AAAI 2026technical

Image forgery localization aims to precisely identify tampered regions within images, but it commonly depends on costly pixel-level annotations. To alleviate this annotation burden, weakly supervised image forgery localization (WSIFL) has emerged, yet existing methods still achieve limited localiza

Cited by 0SourcePDFScholar
2026

Zero-shot Detection of AI-Generated Image via RAW-RGB Alignment

CVPR 2026

Advances in generative AI (GenAI) have increasingly complicated the identification of synthetic images, prompting the proposal of numerous zero-/few-shot detection methods to counter unknown GenAI better. However, we observe that existing detectors often misclassify synthetic images with physical tr

Cited by 0SourceScholar
2025

ADCD-Net: Robust Document Image Forgery Localization via Adaptive DCT Feature and Hierarchical Content Disentanglement

ICCV 2025poster

The advancement of image editing tools has enabled malicious manipulation of sensitive document images, underscoring the need for robust document image forgery detection. Though forgery detectors for natural images have been extensively studied, they struggle with document images, as the tampered re…

2025

An End-to-End Model for Logits-Based Large Language Models Watermarking

ICML 2025poster

The rise of LLMs has increased concerns over source tracing and copyright protection for AIGC, highlighting the need for advanced detection technologies. Passive detection methods usually face high false positives, while active watermarking techniques using logits or sampling manipulation offer more…

2025

Anti-Diffusion: Preventing Abuse of Modifications of Diffusion-Based Models

AAAI 2025technical

Although diffusion-based techniques have shown remarkable success in image generation and editing tasks, their abuse can lead to severe negative social impacts. Recently, some works have been proposed to provide defense against the abuse of diffusion-based methods. However, their protection may be l…

2025

GLCF: A Global-Local Multimodal Coherence Analysis Framework for Talking Face Generation Detection

AAAI 2025technical

Talking face generation (TFG) allows for producing lifelike talking videos of any character using only facial images and accompanying text. Abuse of this technology could pose significant risks to society, creating the urgent need for research into corresponding detection methods. However, research…

Cited by 0SourcePDFScholar
2025

KaRF: Weakly-Supervised Kolmogorov-Arnold Networks-based Radiance Fields for Local Color Editing

NeurIPS 2025poster

Recent advancements have suggested that neural radiance fields (NeRFs) show great potential in color editing within the 3D domain. However, most existing NeRF-based editing methods continue to face significant challenges in local region editing, which usually lead to imprecise local object boundarie…

Cited by 0SourcecodeScholar
2025

OmniMark: Efficient and Scalable Latent Diffusion Model Fingerprinting

AAAI 2025technical

We introduce OmniMark, a novel and efficient fingerprinting method for Latent Diffusion Models (LDM). OmniMark can encode user-specific fingerprints across diverse dimensions of the weights of the LDM, including kernels, filters, channels, and spatial domains. The LDM is fine-tuned to encode the inv…

2025

RaCMC: Residual-Aware Compensation Network with Multi-Granularity Constraints for Fake News Detection

AAAI 2025technical

Multimodal fake news detection aims to automatically identify real or fake news, thereby mitigating the adverse effects caused by such misinformation. Although prevailing approaches have demonstrated their effectiveness, challenges persist in cross-modal feature fusion and refinement for classificat…

Cited by 2SourcePDFScholar
2025

SUMI-IFL: An Information-Theoretic Framework for Image Forgery Localization with Sufficiency and Minimality Constraints

AAAI 2025technical

Image forgery localization (IFL) is a crucial technique for preventing tampered image misuse and protecting social safety. However, due to the rapid development of image tampering technologies, extracting more comprehensive and accurate forgery clues remains an urgent challenge. To address these cha…

Cited by 1SourcePDFScholar
2025

Scalable Dual Fingerprinting for Hierarchical Attribution of Text-to-Image Models

ICCV 2025poster

The commercialization of generative artificial intelligence (GenAI) has led to a multi-level ecosystem involving model developers, service providers, and consumers. Thus, ensuring traceability is crucial, as service providers may violate intellectual property rights (IPR), and consumers may generate…

2025

TurboFill: Adapting Few-step Text-to-image Model for Fast Image Inpainting

CVPR 2025poster

This paper introduces TurboFill, a fast image inpainting model that enhances a few-step text-to-image diffusion model with an inpainting adapter for high-quality and efficient inpainting. While standard diffusion models generate high-quality results, they incur high computational costs. We overcome…

2024

A Unified Environmental Network for Pedestrian Trajectory Prediction

AAAI 2024technical

Accurately predicting pedestrian movements in complex environments is challenging due to social interactions, scene constraints, and pedestrians' multimodal behaviors. Sequential models like long short-term memory fail to effectively integrate scene features to make predicted trajectories comply wit…

Cited by 3SourcePDFScholar
2024

DAT: Improving Adversarial Robustness via Generative Amplitude Mix-up in Frequency Domain

NeurIPS 2024poster

To protect deep neural networks (DNNs) from adversarial attacks, adversarial training (AT) is developed by incorporating adversarial examples (AEs) into model training. Recent studies show that adversarial attacks disproportionately impact the patterns within the phase of the sample's frequency spec…

2024

DifAttack: Query-Efficient Black-Box Adversarial Attack via Disentangled Feature Space

AAAI 2024technical

This work investigates efficient score-based black-box adversarial attacks with high Attack Success Rate (ASR) and good generalizability. We design a novel attack method based on a Disentangled Feature space, called DifAttack, which differs significantly from the existing ones operating over the ent…

2024

Direction-Aware Video Demoiréing with Temporal-Guided Bilateral Learning

AAAI 2024technical

Moiré patterns occur when capturing images or videos on screens, severely degrading the quality of the captured images or videos. Despite the recent progresses, existing video demoiréing methods neglect the physical characteristics and formation process of moiré patterns, significantly limiting the…

Cited by 10SourcePDFScholar
2024

Progressive Poisoned Data Isolation for Training-Time Backdoor Defense

AAAI 2024technical

Deep Neural Networks (DNN) are susceptible to backdoor attacks where malicious attackers manipulate the model's predictions via data poisoning. It is hence imperative to develop a strategy for training a clean model using a potentially poisoned dataset. Previous training-time defense mechanisms typi…

2024

SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models

CVPR 2024highlight

Current instruction-based image editing methods such as InstructPix2Pix often fail to produce satisfactory results in complex scenarios due to their dependence on the simple CLIP text encoder in diffusion models. To rectify this this paper introduces SmartEdit a novel approach of instruction-based i…

2024

Unifying Image Processing as Visual Prompting Question Answering

ICML 2024poster

Image processing is a fundamental task in computer vision, which aims at enhancing image quality and extracting essential features for subsequent vision applications. Traditionally, task-specific models are developed for individual tasks and designing such models requires distinct expertise. Buildin…

2023

Activating More Pixels in Image Super-Resolution Transformer

CVPR 2023poster

Transformer-based methods have shown impressive performance in low-level vision tasks, such as image super-resolution. However, we find that these networks can only utilize a limited spatial range of input information through attribution analysis. This implies that the potential of Transformer is st…

2023

DeSRA: Detect and Delete the Artifacts of GAN-based Real-World Super-Resolution Models

ICML 2023poster

Image super-resolution (SR) with generative adversarial networks (GAN) has achieved great success in restoring realistic details. However, it is notorious that GAN-based SR models will inevitably produce unpleasant and undesirable artifacts, especially in practical scenarios. Previous works typicall…

2023

Effective Ambiguity Attack Against Passport-Based DNN Intellectual Property Protection Schemes Through Fully Connected Layer Substitution

CVPR 2023poster

Since training a deep neural network (DNN) is costly, the well-trained deep models can be regarded as valuable intellectual property (IP) assets. The IP protection associated with deep models has been receiving increasing attentions in recent years. Passport-based method, which replaces normalizatio…

Cited by 16SourcePDFScholar
2023

Hyperspectral Image Denoising Via Nonlocal Rank Residual Modeling

ICASSP 2023accepted

Nonlocal low-rank (LR) tensor modeling has shown great potential in hyperspectral image (HSI) denoising, which first uses the nonlocal self-similarity (NSS) prior to search for many similar full-band patches to form three-dimensional nonlocal full-band groups (tensors), and then usually enforces an…

Cited by 0SourceScholar
2023

Image Sharing Chain Detection VIA Sequence-To-Sequence Model

ICASSP 2023accepted

Image sharing chain detection aims to recover the sharing history of an image downloaded from online social networks (OSNs), including the ever-shared OSNs and their orders, which is an important task in the multimedia forensics community. Most of the existing algorithms directly treat the sharing c…

Cited by 0SourceScholar
2022

Mitigating Inconsistencies in Multimodal Sentiment Analysis under Uncertain Missing Modalities

EMNLP 2022main

For the missing modality problem in Multimodal Sentiment Analysis (MSA), the inconsistency phenomenon occurs when the sentiment changes due to the absence of a modality. The absent modality that determines the overall semantic can be considered as a key missing modality. However, previous works all…

2022

Robust Image Forgery Detection Over Online Social Network Shared Images

CVPR 2022oral

The increasing abuse of image editing softwares, such as Photoshop and Meitu, causes the authenticity of digital images questionable. Meanwhile, the widespread availability of online social networks (OSNs) makes them the dominant channels for transmitting forged images to report fake news, propagate…

Cited by 85PDFcodeScholar
2022

Simultaneous Nonlocal Low-Rank And Deep Priors For Poisson Denoising

ICASSP 2022accepted

Poisson noise is a common electronic noise, which has widely occurred in various photo-limited imaging systems. However, due to signal-dependent and multiplicative characteristics for Poisson noise, Poisson denoising is still an open problem. In this paper, we propose a novel approach using simultan…

Cited by 0SourceScholar
2021

Detecting Adversarial Examples from Sensitivity Inconsistency of Spatial-Transform Domain

AAAI 2021technical

Deep neural networks (DNNs) have been shown to be vulnerable against adversarial examples (AEs), which are maliciously designed to cause dramatic model output errors. In this work, we reveal that normal examples (NEs) are insensitive to the fluctuations occurring at the highly-curved region of the d…

Cited by 67SourcePDFScholar
2021

Probabilistic Selective Encryption of Convolutional Neural Networks for Hierarchical Services

CVPR 2021poster

Model protection is vital when deploying Convolutional Neural Networks (CNNs) for commercial services, due to the massive costs of training them. In this work, we propose a selective encryption (SE) algorithm to protect CNN models from unauthorized access, with a unique feature of providing hierarch…

Cited by 16PDFScholar
2021

Temporal Pyramid Network for Pedestrian Trajectory Prediction with Multi-Supervision

AAAI 2021technical

Predicting human motion behavior in a crowd is important for many applications, ranging from the natural navigation of autonomous vehicles to intelligent security systems of video surveillance. All the previous works model and predict the trajectory with a single resolution, which is relatively inef…

2019

Robust Subspace Clustering With Independent and Piecewise Identically Distributed Noise Modeling

CVPR 2019oral

Most of the existing subspace clustering (SC) frameworks assume that the noise contaminating the data is generated by an independent and identically distributed (i.i.d.) source, where the Gaussianity is often imposed. Though these assumptions greatly simplify the underlying problems, they do not hol…

Cited by 22PDFScholar
2015

Data-Driven Sparsity-Based Restoration of JPEG-Compressed Images in Dual Transform-Pixel Domain

CVPR 2015poster

Arguably the most common cause of image degradation is compression. This papers presents a novel approach to restoring JPEG-compressed images. The main innovation is in the approach of exploiting residual redundancies of JPEG code streams and sparsity properties of latent images. The restoration…

Cited by 104SourcePDFScholar