← Search

Jae Hyeon Cho

3 accepted papers

2026

Beyond RAG vs. Long-Context: Learning Distraction-Aware Retrieval for Efficient Knowledge Grounding

ICLR 2026poster

Retrieval-Augmented Generation (RAG) is a framework for grounding Large Language Models (LLMs) in external, up-to-date information. However, recent advancements in context window size allow LLMs to process inputs of up to 128K tokens or more, offering an alternative strategy: supplying the full docu…

Cited by 0SourceScholar
2025

K/DA: Automated Data Generation Pipeline for Detoxifying Implicitly Offensive Language in Korean

ACL 2025long

Language detoxification involves removing toxicity from offensive language. While a neutral-toxic paired dataset provides a straightforward approach for training detoxification models, creating such datasets presents several challenges: i) the need for human annotation to build paired data, and ii)…

2025

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment

EMNLP 2025

Direct Preference Optimization (DPO) is a simple and efficient framework that has attracted substantial attention. However, it often struggles to meet its primary objectives—increasing the generation probability of chosen responses while reducing that of rejected responses—due to the dominant influe