← Search

Heuiyeen Yeen

2 accepted papers

2025

Ko-LongRAG: A Korean Long-Context RAG Benchmark Built with a Retrieval-Free Approach

EMNLP 2025

The rapid advancement of large language models (LLMs) significantly enhances long-context Retrieval-Augmented Generation (RAG), yet existing benchmarks focus primarily on English. This leaves low-resource languages without comprehensive evaluation frameworks, limiting their progress in retrieval-bas

2025

MANTA: A Scalable Pipeline for Transmuting Massive Web Corpora into Instruction Datasets

EMNLP 2025

We introduce MANTA, an automated pipeline that generates high-quality large-scale instruction fine-tuning datasets from massive web corpora while preserving their diversity and scalability. By extracting structured syllabi from web documents and leveraging high-performance LLMs, our approach enables