Temper-Then-Tilt: Principled Unlearning for Generative Models through Tempering and Classifier Guidance
We study machine unlearning in large generative models by framing the task as density ratio estimation to a target distribution rather than supervised fine-tuning. While classifier guidance is a standard approach for approximating this ratio and can succeed in general, we show it can fail to faithfu…