← Search

Ramon Huerta

1 accepted papers

2025

UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models

NAACL 2025long

Mitigating the retention of sensitive or private information in large language models is essential for enhancing privacy and safety. Existing unlearning methods, like Gradient Ascent and Negative Preference Optimization, directly tune models to remove unwanted information. However, these methods oft…