← Search

Hyunwoo Ko

3 accepted papers

2026

Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math

ICML 2026spotlight

Recent progress in reasoning models suggests that generating plausible attempts for research-level mathematics may be within reach, but verification remains a bottleneck, consuming scarce expert time. We hypothesize that a meaningful solution should contain enough method-level information that, when…

Cited by 0SourceScholar
2026

Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought

ICLR 2026poster

Recent frontier models employ long-chain-of-thought reasoning to explore solution spaces in context and achieve stronger performance. While many works study distillation to build smaller yet capable models, most focus on English and little is known about language-specific reasoning. To bridge this g…

Cited by 0SourcecodeScholar
2025

Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning

ACL 2025long

Scaling pre-training compute has proven effective for achieving multilinguality, but does the same hold for test-time scaling? In this work, we introduce **MCLM**, a multilingual math benchmark featuring competition-level problems in 55 languages. We then compare three test-time scaling methods—Outc…