← Search

Nay Myat Min

2 accepted papers

2026

Propaganda AI: An Analysis of Semantic Divergence in Large Language Models

ICLR 2026poster

Large language models (LLMs) can exhibit *concept-conditioned semantic divergence*: common high-level cues (e.g., ideologies, public figures) elicit unusually uniform, stance-like responses that evade token-trigger audits. This behavior falls in a blind spot of current safety evaluations, yet carrie…

Cited by 0SourcecodeScholar
2025

CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization

ICML 2025poster

Large Language Models (LLMs) are vulnerable to backdoor attacks that manipulate outputs via hidden triggers. Existing defense methods—designed for vision/text classification tasks—fail for text generation. We propose *Internal Consistency Regularization (CROW)*, a defense leveraging the observation…