← Search

Jay Shin

3 accepted papers

2026

Predicting LLM Reasoning Performance with Small Proxy Model

ICLR 2026poster

Given the prohibitive cost of pre-training large language models, it is essential to leverage smaller proxy models to optimize recipes before scaling up. However, this approach becomes challenging for reasoning capabilities, which exhibit \textit{emergent} behavior that only appears reliably at larg…

Cited by 0SourcecodeScholar
2025

MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language

EMNLP 2025

Evaluating text generation capabilities of large language models (LLMs) is challenging, particularly for low-resource languages where methods for direct assessment are scarce. We propose MUG-Eval, a novel framework that evaluates LLMs’ multilingual generation capabilities by transforming existing be