← Search

Moayad Aloqaily

3 accepted papers

2026

AutoDebias: An Automated Framework for Detecting and Mitigating Backdoor Biases in Text-to-Image Models

CVPR 2026

Text-to-Image (T2I) models generate high-quality images but are vulnerable to malicious backdoor attacks that inject harmful biases (e.g., trigger-activated gender or racial stereotypes). Existing debiasing methods, often designed for natural statistical biases, struggle with these deliberate and su

Cited by 0SourcecodeScholar
2026

SafeSeek: Universal Attribution of Safety Circuits in Language Models

ICML 2026poster

Mechanistic interpretability reveals that safety-critical behaviors (e.g., alignment, jailbreak, backdoor) in Large Language Models (LLMs) are grounded in specialized functional components. However, existing safety attribution methods struggle with generalization and reliability due to their relianc…

Cited by 0SourceScholar
2026

Uncovering Hidden Triggers: Backdoor Attribution in Language Models

ICML 2026poster

Fine-tuned Large Language Models (LLMs) are vulnerable to backdoor attacks through data poisoning, yet the internal mechanisms governing these attacks remain a black box. Previous research on interpretability for LLM safety tends to focus on alignment, jailbreak, and hallucination, but overlooks bac…

Cited by 0SourceScholar