← Search

Kyohoon Jin

7 accepted papers

2025

CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples

EMNLP 2025

Deep learning models often learn and exploit spurious correlations in training data, using these non-target features to inform their predictions. Such reliance leads to performance degradation and poor generalization on unseen data. To address these limitations, we introduce a more general form of c

Cited by 0SourcePDFScholar
2025

GRADE: Generating multi-hop QA and fine-gRAined Difficulty matrix for RAG Evaluation

EMNLP 2025

Retrieval-Augmented Generation (RAG) systems are widely adopted in knowledge-intensive NLP tasks, but current evaluations often overlook the structural complexity and multi-step reasoning required in real-world scenarios. These benchmarks overlook key factors such as the interaction between retrieva

2025

Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models

ACL 2025long

Large language models (LLMs) are renowned for their extensive linguistic knowledge and strong generalization capabilities, but their high computational demands make them unsuitable for resource-constrained environments. In contrast, small language models (SLMs) are computationally efficient but ofte…

2025

SummPilot: Bridging Efficiency and Customization for Interactive Summarization System

AAAI 2025technical

This paper incorporates the efficiency of automatic summarization and addresses the challenge of generating personalized summaries tailored to individual users' interests and requirements. To tackle this challenge, we introduce SummPilot, an interaction-based customizable summarization system. SummP…

Cited by 0SourcePDFScholar
2024

Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation

COLING 2024main

Efforts to leverage deep learning models in low-resource regimes have led to numerous augmentation studies. However, the direct application of methods, such as mixup and cutout, is limited due to the discrete characteristics of the textual data. While methods using pre trained language models have e…

Cited by 0SourcePDFScholar
2024

Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation

EMNLP 2024main

The quality of the dataset is crucial for ensuring optimal performance and reliability of downstream task models. However, datasets often contain noisy data inadvertently included during the construction process. Numerous attempts have been made to correct this issue through human annotators. Howeve…

2021

Restoring and Mining the Records of the Joseon Dynasty via Neural Language Modeling and Machine Translation

NAACL 2021long

Understanding voluminous historical records provides clues on the past in various aspects, such as social and political issues and even natural science facts. However, it is generally difficult to fully utilize the historical records, since most of the documents are not written in a modern language…