← Search

Seungyoon Lee

8 accepted papers

2025

FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models

NAACL 2025findings

Recent advancements in Large Language Models (LLMs) have significantly enhanced interactions between users and models. These advancements concurrently underscore the need for rigorous safety evaluations due to the manifestation of social biases, which can lead to harmful societal impacts. Despite th…

2025

Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models

NAACL 2025findings

Through numerous endeavors, large language models (LLMs) have witnessed significant advancements in their instruction-following capability. However, we discern that LLMs are prone to generate responses to instruction-formatted statements in an instinctive manner, rather than comprehending the underl…

2025

MIGRATE: Cross-Lingual Adaptation of Domain-Specific LLMs through Code-Switching and Embedding Transfer

COLING 2025main

Large Language Models (LLMs) have rapidly advanced, with domain-specific expert models emerging to handle specialized tasks across various fields. However, the predominant focus on English-centric models demands extensive data, making it challenging to develop comparable models for middle and low-re…

Cited by 0SourcePDFScholar
2025

Semantic Aware Linear Transfer by Recycling Pre-trained Language Models for Cross-lingual Transfer

ACL 2025finding

Large Language Models (LLMs) are increasingly incorporating multilingual capabilities, fueling the demand to transfer them into target language-specific models. However, most approaches, which blend the source model’s embedding by replacing the source vocabulary with the target language-specific voc…

Cited by 0SourcePDFScholar
2024

Length-aware Byte Pair Encoding for Mitigating Over-segmentation in Korean Machine Translation

ACL 2024findings

Byte Pair Encoding is an effective approach in machine translation across several languages. However, our analysis indicates that BPE is prone to over-segmentation in the morphologically rich language, Korean, which can erode word semantics and lead to semantic confusion during training. This semant…

2024

Leveraging Pre-existing Resources for Data-Efficient Counter-Narrative Generation in Korean

COLING 2024main

Counter-narrative generation, i.e., the generation of fact-based responses to hate speech with the aim of correcting discriminatory beliefs, has been demonstrated to be an effective method to combat hate speech. However, its effectiveness is limited by the resource-intensive nature of dataset constr…

2024

Translation of Multifaceted Data without Re-Training of Machine Translation Systems

EMNLP 2024finding

Translating major language resources to build minor language resources becomes a widely-used approach. Particularly in translating complex data points composed of multiple components, it is common to translate each component separately. However, we argue that this practice often overlooks the interr…

Cited by 0SourcePDFScholar