2024
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
ICLR 2024spotlight
Modern language models can imitate complex patterns through few-shot learning, enabling them to complete challenging tasks without fine-tuning. However, imitation can also lead models to reproduce inaccuracies or harmful content if present in the context. We study harmful imitation through the lens…