← Search

Hyeonseok Lim

5 accepted papers

2025

Can LLMs Truly Plan? A Comprehensive Evaluation of Planning Capabilities

EMNLP 2025

The existing assessments of planning capabilities of large language models (LLMs) remain largely limited to single-language or specific representation formats. To address this gap, we introduce the Multi-Plan benchmark comprising 204 multilingual and multi-format travel planning scenarios. In experi

Cited by 0SourcePDFScholar
2025

Unified Automated Essay Scoring and Grammatical Error Correction

NAACL 2025findings

This study explores the integration of automated writing evaluation (AWE) and grammatical error correction (GEC) through multitask learning, demonstrating how combining these distinct tasks can enhance performance in both areas. By leveraging a shared learning framework, we show that models trained…

Cited by 1SourcePDFScholar
2025

VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation

COLING 2025main

We propose the VLR-Bench, a visual question answering (VQA) benchmark for evaluating vision language models (VLMs) based on retrieval augmented generation (RAG). Unlike existing evaluation datasets for external knowledge-based VQA, the proposed VLR-Bench includes five input passages. This allows tes…

Cited by 1SourcePDFScholar
2024

Optimizing Language Augmentation for Multilingual Large Language Models: A Case Study on Korean

COLING 2024main

Large language models (LLMs) use pretraining to predict the subsequent word; however, their expansion requires significant computing resources. Numerous big tech companies and research institutes have developed multilingual LLMs (MLLMs) to meet current demands, overlooking less-resourced languages (…

2024

X-LLaVA: Optimizing Bilingual Large Vision-Language Alignment

NAACL 2024findings

The impressive development of large language models (LLMs) is expanding into the realm of large multimodal models (LMMs), which incorporate multiple types of data beyond text. However, the nature of multimodal models leads to significant expenses in the creation of training data. Furthermore, constr…