← Search

Shagufta Mehnaz

3 accepted papers

2026

Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models

AAAI 2026technical

Large Language Models (LLMs) exhibit impressive capabilities, but remain susceptible to a growing spectrum of safety risks, including jailbreaks, toxic content, hallucinations, and bias. Existing defenses often address only a single threat type or resort to rigid outright rejection, sacrificing user

Cited by 0SourcePDFScholar
2025

Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage

AAAI 2025technical

Fine-tuning large language models on private data for downstream applications poses significant privacy risks in potentially exposing sensitive information. Several popular community platforms now offer convenient distribution of a large variety of pre-trained models, allowing anyone to publish with…

Cited by 4SourcePDFScholar
2025

From Insight to Exploit: Leveraging LLM Collaboration for Adaptive Adversarial Text Generation

EMNLP 2025

LLMs can provide substantial zero-shot performance on diverse tasks using a simple task prompt, eliminating the need for training or fine-tuning. However, when applying these models to sensitive tasks, it is crucial to thoroughly assess their robustness against adversarial inputs. In this work, we i