← Search

Hyukhun Koh

12 accepted papers

2026

Confidence-Guided Stepwise Model Routing for Cost-Efficient Reasoning

AAAI 2026technical

Recent advances in Large Language Models (LLMs) - particularly model scaling and test-time techniques - have greatly enhanced the reasoning capabilities of language models at the expense of higher inference costs. To lower inference costs, prior works train router models or deferral mechanisms that

Cited by 0SourcePDFScholar
2025

Generating Diverse Hypotheses for Inductive Reasoning

NAACL 2025long

Inductive reasoning — the process of inferring general rules from a small number of observations — is a fundamental aspect of human intelligence. Recent works suggest that large language models (LLMs) can engage in inductive reasoning by sampling multiple hypotheses about the rules and selecting the…

Cited by 0SourcePDFScholar
2025

Harmful Prompt Laundering: Jailbreaking LLMs with Abductive Styles and Symbolic Encoding

EMNLP 2025

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but their potential misuse for harmful purposes remains a significant concern. To strengthen defenses against such vulnerabilities, it is essential to investigate universal jailbreak attacks that exploit int

Cited by 0SourcePDFScholar
2025

Program Synthesis via Test-Time Transduction

NeurIPS 2025poster

We introduce transductive program synthesis, a new formulation of the program synthesis task that explicitly leverages test inputs during synthesis. While prior approaches to program synthesis--whether based on natural language descriptions or input-output examples--typically aim to generalize from…

Cited by 2SourcecodeScholar
2025

VLind-Bench: Measuring Language Priors in Large Vision-Language Models

NAACL 2025findings

Large Vision-Language Models (LVLMs) have demonstrated outstanding performance across various multimodal tasks. However, they suffer from a problem known as language prior, where responses are generated based solely on textual patterns while disregarding image information. Addressing the issue of la…

2024

Can LLMs Recognize Toxicity? A Structured Investigation Framework and Toxicity Metric

EMNLP 2024finding

In the pursuit of developing Large Language Models (LLMs) that adhere to societal standards, it is imperative to detect the toxicity in the generated text. The majority of existing toxicity metrics rely on encoder models trained on specific toxicity datasets, which are susceptible to out-of-distribu…

2024

Fine-grained Gender Control in Machine Translation with Large Language Models

NAACL 2024long

In machine translation, the problem of ambiguously gendered input has been pointed out, where the gender of an entity is not available in the source sentence. To address this ambiguity issue, the task of controlled translation that takes the gender of the ambiguous entity as additional input have be…

Cited by 3SourcePDFScholar
2023

DPP-TTS: Diversifying prosodic features of speech via determinantal point processes

EMNLP 2023long main

With the rapid advancement in deep generative models, recent neural Text-To-Speech(TTS) models have succeeded in synthesizing human-like speech. There have been some efforts to generate speech with various prosody beyond monotonous prosody patterns. However, previous works have several limitations.…

Cited by 0SourceScholar
2023

Target-Agnostic Gender-Aware Contrastive Learning for Mitigating Bias in Multilingual Machine Translation

EMNLP 2023long main

Gender bias is a significant issue in machine translation, leading to ongoing research efforts in developing bias mitigation techniques. However, most works focus on debiasing bilingual models without much consideration for multilingual systems. In this paper, we specifically target the gender bias…

Cited by 0SourcecodeScholar