← Search

Jinta Weng

2 accepted papers

2026

Steering Representations, Safeguarding Privacy: A Cross-Modal Privacy Protection Method for Generative AI

AAAI 2026technical

Privacy concerns have long been a critical issue in AI models. With the rapid advancement of generative AI, the privacy awareness of models has drawn attention, raising new challenges for privacy protection that is independent of data and tasks. This paper introduces a novel framework for enhancing

Cited by 0SourcePDFScholar
2025

RepGuard: Adaptive Feature Decoupling for Robust Backdoor Defense in Large Language Models

NeurIPS 2025poster

Backdoor attacks pose a significant threat to large language models (LLMs) by embedding malicious triggers that manipulate model behavior. However, existing defenses primarily rely on prior knowledge of backdoor triggers or targets and offer only superficial mitigation strategies, thus struggling to…

Cited by 0SourceScholar