← Search

Zhangjie Fu

6 accepted papers

2026

Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration

CVPR 2026

Text-to-image generative models have achieved impressive fidelity and diversity, but can inadvertently produce unsafe or undesirable content due to implicit biases embedded in large-scale training datasets.Existing concept erasure methods, whether text-only or image-assisted, face trade-offs: textua

Cited by 0SourceScholar
2026

DGS-Net: Distillation-Guided Gradient Surgery for CLIP Fine-Tuning in AI-Generated Image Detection

ICML 2026spotlight

The rapid progress of generative models such as GANs and diffusion models has led to the widespread proliferation of AI-generated images, raising concerns about misinformation, privacy violations, and trust erosion in digital media. Although large-scale multimodal models like CLIP offer strong trans…

Cited by 0SourceScholar
2026

No Way To Steal My Face: Proactive Defense Against Identity-Preserving Personalized Generation

CVPR 2026

Recent advances in diffusion models have enabled high-fidelity, identity-preserving image generation for personalized applications such as digital avatars and virtual try-on systems. However, their reliance on sensitive facial reference images raises growing privacy concerns. Existing defense mechan

Cited by 0SourceScholar
2025

Is Artificial Intelligence Generated Image Detection a Solved Problem?

NeurIPS 2025poster

The rapid advancement of generative models, such as GANs and Diffusion models, has enabled the creation of highly realistic synthetic images, raising serious concerns about misinformation, deepfakes, and copyright infringement. Although numerous Artificial Intelligence Generated Image (AIGI) detecto…

Cited by 0SourcecodeScholar
2023

CurveFormer: 3D Lane Detection by Curve Propagation with Curve Queries and Attention

ICRA 2023poster

3D lane detection is an integral part of au-tonomous driving systems. Previous CNN and Transformer-based methods usually first generate a bird's-eye-view (BEV) feature map from the front view image, and then use a sub-network with BEV feature map as input to predict 3D lanes. Such approaches require…

Cited by 60SourceScholar