← Search

Junghun Yuk

3 accepted papers

2025

ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts

EMNLP 2025

Prior benchmarks for evaluating the domain-specific knowledge of large language models (LLMs) lack the scalability to handle complex academic tasks. To address this, we introduce ScholarBench, a benchmark centered on deep expert knowledge and complex academic problem-solving, which evaluates the aca

Cited by 1SourcePDFScholar
2025

Unified Automated Essay Scoring and Grammatical Error Correction

NAACL 2025findings

This study explores the integration of automated writing evaluation (AWE) and grammatical error correction (GEC) through multitask learning, demonstrating how combining these distinct tasks can enhance performance in both areas. By leveraging a shared learning framework, we show that models trained…

Cited by 1SourcePDFScholar
2025

VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation

COLING 2025main

We propose the VLR-Bench, a visual question answering (VQA) benchmark for evaluating vision language models (VLMs) based on retrieval augmented generation (RAG). Unlike existing evaluation datasets for external knowledge-based VQA, the proposed VLR-Bench includes five input passages. This allows tes…

Cited by 1SourcePDFScholar