← Search

Nakyeong Yang

7 accepted papers

2026

Bilinear relational structure fixes reversal curse and enables consistent model editing

ICLR 2026poster

The reversal curse---a language model's (LM) inability to infer an unseen fact ``B is A'' from a learned factA is B''---is widely considered a fundamental limitation. We show that this is not an inherent failure but an artifact of how models encode knowledge. By training LMs from scratch on a synthe…

Cited by 0SourceScholar
2026

Confidence-Guided Stepwise Model Routing for Cost-Efficient Reasoning

AAAI 2026technical

Recent advances in Large Language Models (LLMs) - particularly model scaling and test-time techniques - have greatly enhanced the reasoning capabilities of language models at the expense of higher inference costs. To lower inference costs, prior works train router models or deferral mechanisms that

Cited by 0SourcePDFScholar
2026

Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning

ICLR 2026poster

Large language models trained on web-scale data can memorize private or sensitive knowledge, raising significant privacy risks. Although some unlearning methods mitigate these risks, they remain vulnerable to "relearning" during subsequent training, allowing a substantial portion of forgotten knowle…

Cited by 0SourceScholar
2025

FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge

EMNLP 2025

Various studies have attempted to remove sensitive or private knowledge from a language model to prevent its unauthorized exposure. However, prior studies have overlooked the inherent complexity and interconnectedness of knowledge, which requires careful examination. To resolve this problem, we firs

2024

Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination

ACL 2024long

Instruction-following language models often show undesirable biases. These undesirable biases may be accelerated in the real-world usage of language models, where a wide range of instructions is used through zero-shot example prompting. To solve this problem, we first define the bias neuron, which s…

Cited by 7SourcePDFScholar
2022

Deriving Explainable Discriminative Attributes Using Confusion About Counterfactual Class

ICASSP 2022accepted

Recently, Integrated Gradients-based (IG) methods have been commonly used to explain the decision process of deep neural networks (DNNs). However, they have only considered the information of the predicted class while neglecting the in-formation of the rest classes. In this paper, we propose a novel…

Cited by 0SourceScholar