← Search

Sangyeon Yoon

6 accepted papers

2026

A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models

ICLR 2026poster

Diffusion large language models (dLLMs) enable any-order generation, but this flexibility enlarges the attack surface: harmful spans may appear at arbitrary positions, and template-based prefilling attacks such as DIJA bypass response-level refusals. We introduce A2D (Any-Order, Any-Step Defense), a…

Cited by 0SourcecodeScholar
2026

Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures

ICLR 2026poster

Machine unlearning aims to remove specific content from trained models while preserving overall performance. However, the phenomenon of benign relearning, in which forgotten information reemerges even from benign fine-tuning data, reveals that existing unlearning methods remain fundamentally fragile…

Cited by 0SourceScholar
2025

SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment

NeurIPS 2025poster

Large Reasoning Models (LRMs) have become powerful tools for complex problem solving, but their structured reasoning pathways can lead to unsafe outputs when exposed to harmful prompts. Existing safety alignment methods reduce harmful outputs but can degrade reasoning depth, leading to significant t…

Cited by 0SourceScholar