Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
We introduce and study artificial impressions–patterns in LLMs’ internal representations of prompts that resemble human impressions and stereotypes based on language. We fit linear probes on generated prompts to predict impressions according to the two-dimensional Stereotype Content Model (SCM). Usi