← Search

Wonje Jeung

11 accepted papers

2026

A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models

ICLR 2026poster

Diffusion large language models (dLLMs) enable any-order generation, but this flexibility enlarges the attack surface: harmful spans may appear at arbitrary positions, and template-based prefilling attacks such as DIJA bypass response-level refusals. We introduce A2D (Any-Order, Any-Step Defense), a…

Cited by 0SourcecodeScholar
2026

An Information Theoretic Evaluation Metric for Strong Unlearning

AAAI 2026technical

Machine unlearning (MU) aims to remove the influence of specific data from trained models, addressing privacy concerns and ensuring compliance with regulations such as the "right to be forgotten." Evaluating strong unlearning, where the unlearned model is indistinguishable from one retrained without

Cited by 0SourcePDFScholar
2026

Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMs

ICLR 2026poster

Diffusion large language models (dLLMs) have emerged as a promising alternative to autoregressive models, offering flexible generation orders and strong performance on complex reasoning tasks. However, instruction-tuned dLLMs exhibit a critical vulnerability we term \<eos\> overflow: as allocated s…

Cited by 0SourceScholar
2026

Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning Failures

ICLR 2026poster

Machine unlearning aims to remove specific content from trained models while preserving overall performance. However, the phenomenon of benign relearning, in which forgotten information reemerges even from benign fine-tuning data, reveals that existing unlearning methods remain fundamentally fragile…

Cited by 0SourceScholar
2025

Large Language Models Still Exhibit Bias in Long Text

ACL 2025finding

Existing fairness benchmarks for large language models (LLMs) primarily focus on simple tasks, such as multiple-choice questions, overlooking biases that may arise in more complex scenarios like long-text generation. To address this gap, we introduce the Long Text Fairness Test (LTF-TEST), a framewo…

Cited by 0SourcePDFScholar
2025

Representation Bending for Large Language Model Safety

ACL 2025long

Large Language Models (LLMs) have emerged as powerful tools, but their inherent safety risks – ranging from harmful content generation to broader societal harms – pose significant challenges. These risks can be amplified by the recent adversarial attacks, fine-tuning vulnerabilities, and the increas…

2025

SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment

NeurIPS 2025poster

Large Reasoning Models (LRMs) have become powerful tools for complex problem solving, but their structured reasoning pathways can lead to unsafe outputs when exposed to harmful prompts. Existing safety alignment methods reduce harmful outputs but can degrade reasoning depth, leading to significant t…

Cited by 0SourceScholar
2024

Learning Equi-angular Representations for Online Continual Learning

CVPR 2024poster

Online continual learning suffers from an underfitted solution due to insufficient training for prompt model updates (e.g. single-epoch training). To address the challenge we propose an efficient online continual learning method using the neural collapse phenomenon. In particular we induce neural co…

2024

ReALFRED: An Embodied Instruction Following Benchmark in Photo-Realistic Environments

ECCV 2024poster

"Simulated virtual environments have been widely used to learn robotic agents that perform daily household tasks. These environments encourage research progress by far, but often provide limited object interactability, visual appearance different from real-world environments, or relatively smaller e…