← Search

Wenbo Zhou

26 accepted papers

2026

Design of an Active Haptic Interface Using Proprioception Feedback for Continuous Endovascular Teleoperation

RA-L 2026

Force feedback is essential for safe endovascular teleoperation, yet typically constrained by complex sensor integration. This article presents a compact active haptic interface system designed for robotic catheterization. Leveraging the intrinsic proprioception of a Permanent Magnet Synchronous Mot

Cited by 0SourceScholar
2026

EARG-Net: Edge-Aware Reconstruction-Guided Network for Image Manipulation Detection and Localization

AAAI 2026technical

Recent advances in image editing tools, particularly those used in content-aware retouching and object-level manipulation, have raised significant concerns regarding the authenticity of digital images. While many Image Manipulation Detection and Localization (IMDL) methods have been proposed, they o

Cited by 0SourcePDFScholar
2026

MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement

AAAI 2026technical

Generating expressive and controllable human speech is one of the core goals of generative artificial intelligence, but its progress has long been constrained by two fundamental challenges: the deep entanglement of speech factors and the coarse granularity of existing control mechanisms. To overcome

Cited by 0SourcePDFScholar
2026

Real Data Lies: Unveiling and Closing the Quality Shortcut in Generalizable AI-Generated Video Detection

ICML 2026poster

Recent advances in video generation have enabled highly realistic synthetic content, raising concerns about the integrity of digital media and motivating the development of benchmarks and detection methods for generated videos. Prior works have largely prioritized bolstering model generalization aga…

Cited by 0SourceScholar
2026

State-Dependent Safety Failures in Multi-Turn Language Model Interaction

ICML 2026poster

Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although multi-turn jailbreaks are empirically effective, the structure of conversational safety failure remains insufficiently understood. In this work, we study saf…

Cited by 0SourceScholar
2025

CASAGPT: Cuboid Arrangement and Scene Assembly for Interior Design

CVPR 2025highlight

We present a novel approach for indoor scene synthesis, which learns to arrange decomposed cuboid primitives to represent 3D objects within a scene. Unlike conventional methods that use bounding boxes to determine the placement and scale of 3D objects, our approach leverages cuboids as a straightfor…

Cited by 1SourcePDFScholar
2025

Efficient Dataset Distillation through Low-Rank Space Sampling

ICASSP 2025accepted

Huge amount of data is the key of the success of deep learning, however, redundant information impairs the generalization ability of the model and increases the burden of calculation. Dataset Distillation (DD) compresses the original dataset into a smaller but representative subset for high-quality…

Cited by 0SourceScholar
2025

EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers

ICML 2025poster

Removing unwanted concepts from large-scale text-to-image (T2I) diffusion models while maintaining their overall generative quality remains an open challenge. This difficulty is especially pronounced in emerging paradigms, such as Stable Diffusion (SD) v3 and Flux, which incorporate flow matching an…

2025

MES-RAG: Bringing Multi-modal, Entity-Storage, and Secure Enhancements to RAG

NAACL 2025findings

Retrieval-Augmented Generation (RAG) improves Large Language Models (LLMs) by using external knowledge, but it struggles with precise entity information retrieval. Our proposed **MES-RAG** framework enhances entity-specific query handling and provides accurate, secure, and consistent responses. MES-…

2025

Segue: Side-information Guided Generative Unlearnable Examples for Facial Privacy Protection in Real World

ICASSP 2025accepted

The widespread adoption of face recognition has raised privacy concerns regarding the collection and use of facial data. To address this, researchers have explored "unlearnable examples" by adding imperceptible perturbations during model training to prevent the model from learning target features. H…

Cited by 0SourceScholar
2025

T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks

NeurIPS 2025poster

In recent years, fueled by the rapid advancement of diffusion models, text-to-video (T2V) generation models have achieved remarkable progress, with notable examples including Pika, Luma, Kling, and Open-Sora. Although these models exhibit impressive generative capabilities, they also expose signific…

Cited by 0SourceScholar
2024

AquaLoRA: Toward White-box Protection for Customized Stable Diffusion Models via Watermark LoRA

ICML 2024poster

Diffusion models have achieved remarkable success in generating high-quality images. Recently, the open-source models represented by Stable Diffusion (SD) are thriving and are accessible for customization, giving rise to a vibrant community of creators and enthusiasts. However, the widespread availa…

2024

Attribute-Aware Head Swapping Guided by 3d Modeling

ICASSP 2024accepted

Face manipulation has ignited the interests of both academia and industry in very recent years. Existing face manipulation methods can be roughly categorized into two types: face attribute editing and face swapping. In this paper, we focus on swapping the identity. But unlike face swapping which onl…

Cited by 0SourceScholar
2024

FaceRSA: RSA-Aware Facial Identity Cryptography Framework

AAAI 2024technical

With the flourishing of the Internet, sharing one's photos or automated processing of faces using computer vision technology has become an everyday occurrence. While enjoying the convenience, the concern for identity privacy is also emerging. Therefore, some efforts introduced the concept of ``passw…

Cited by 2SourcePDFScholar
2024

Transferable Facial Privacy Protection against Blind Face Restoration via Domain-Consistent Adversarial Obfuscation

ICML 2024poster

With the rise of social media and the proliferation of facial recognition surveillance, concerns surrounding privacy have escalated significantly. While numerous studies have concentrated on safeguarding users against unauthorized face recognition, a new and often overlooked issue has emerged due to…

Cited by 1SourcePDFScholar
2023

HairCLIPv2: Unifying Hair Editing via Proxy Feature Blending

ICCV 2023poster

Hair editing has made tremendous progress in recent years. Early hair editing methods use well-drawn sketches or masks to specify the editing conditions. Even though they can enable very fine-grained local control, such interaction modes are inefficient for the editing conditions that can be easily…

Cited by 22PDFcodeScholar
2023

X-Paste: Revisiting Scalable Copy-Paste for Instance Segmentation using CLIP and StableDiffusion

ICML 2023poster

Copy-Paste is a simple and effective data augmentation strategy for instance segmentation. By randomly pasting object instances onto new background images, it creates new training data for free and significantly boosts the segmentation performance, especially for rare object categories. Although div…

2022

FInfer: Frame Inference-Based Deepfake Detection for High-Visual-Quality Videos

AAAI 2022technical

Deepfake has ignited hot research interests in both academia and industry due to its potential security threats. Many countermeasures have been proposed to mitigate such risks. Current Deepfake detection methods achieve superior performances in dealing with low-visual-quality Deepfake media which ca…

2022

HairCLIP: Design Your Hair by Text and Reference Image

CVPR 2022poster

Hair editing is an interesting and challenging problem in computer vision and graphics. Many existing methods require well-drawn sketches or masks as conditional inputs for editing, however these interactions are neither straightforward nor efficient. In order to free users from the tedious interact…

Cited by 133PDFcodeScholar
2021

Improved Image Matting via Real-Time User Clicks and Uncertainty Estimation

CVPR 2021poster

Image matting is a fundamental and challenging problem in computer vision and graphics. Most existing matting methods leverage a user-supplied trimap as an auxiliary input to produce good alpha matte. However, obtaining high-quality trimap itself is arduous, thus restricting the application of these…

Cited by 41PDFScholar
2021

Initiative Defense against Facial Manipulation

AAAI 2021technical

Benefiting from the development of generative adversarial networks (GAN), facial manipulation has achieved significant progress in both academia and industry recently. It inspires an increasing number of entertainment applications but also incurs severe threats to individual privacy and even politic…

2021

Multi-Attentional Deepfake Detection

CVPR 2021poster

Face forgery by deepfake is widely spread over the internet and has raised severe societal concerns. Recently, how to detect such forgery contents has become a hot research topic and many deepfake detection methods have been proposed. Most of them model deepfake detection as a vanilla binary classif…

Cited by 885PDFcodeScholar
2021

Spatial-Phase Shallow Learning: Rethinking Face Forgery Detection in Frequency Domain

CVPR 2021poster

The remarkable success in face forgery techniques has received considerable attention in computer vision due to security concerns. We observe that up-sampling is a necessary step of most face forgery techniques, and cumulative up-sampling will result in obvious changes in the frequency domain, espec…

Cited by 517PDFScholar
2019

DUP-Net: Denoiser and Upsampler Network for 3D Adversarial Point Clouds Defense

ICCV 2019poster

Neural networks are vulnerable to adversarial examples, which poses a threat to their application in security sensitive systems. We propose a Denoiser and UPsampler Network (DUP-Net) structure as defenses for 3D adversarial point cloud classification, where the two modules reconstruct surface smooth…

Cited by 202PDFcodeScholar