2026
Propaganda AI: An Analysis of Semantic Divergence in Large Language Models
ICLR 2026poster
Large language models (LLMs) can exhibit *concept-conditioned semantic divergence*: common high-level cues (e.g., ideologies, public figures) elicit unusually uniform, stance-like responses that evade token-trigger audits. This behavior falls in a blind spot of current safety evaluations, yet carrie…