← Search

Jiwoo Hong

10 accepted papers

2026

Margin-Aware Preference Optimization for Aligning Diffusion Models Without Reference

AAAI 2026technical

Modern preference alignment methods, such as DPO, rely on divergence regularization to a reference model for training stability—but this creates a fundamental problem we call "reference mismatch." In this paper, we investigate the negative impacts of reference mismatch in aligning text-to-image (T2I

Cited by 0SourcePDFScholar
2025

AlphaPO: Reward Shape Matters for LLM Alignment

ICML 2025poster

Reinforcement Learning with Human Feedback (RLHF) and its variants have made huge strides toward the effective alignment of large language models (LLMs) to follow instructions and reflect human values. More recently, Direct Alignment Algorithms (DAAs) have emerged in which the reward modeling stage…

Cited by 0SourcePDFScholar
2025

Cross-lingual Transfer of Reward Models in Multilingual Alignment

NAACL 2025short

Reinforcement learning with human feedback (RLHF) is shown to largely benefit from precise reward models (RMs). However, recent studies in reward modeling schemes are skewed towards English, limiting the applicability of RLHF in multilingual alignments. In this work, we investigate the cross-lingual…

2025

Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning

ACL 2025long

Scaling pre-training compute has proven effective for achieving multilinguality, but does the same hold for test-time scaling? In this work, we introduce **MCLM**, a multilingual math benchmark featuring competition-level problems in 55 languages. We then compare three test-time scaling methods—Outc…

2025

On the Robustness of Reward Models for Language Model Alignment

ICML 2025poster

The Bradley-Terry (BT) model is widely practiced in reward modeling for reinforcement learning with human feedback (RLHF). Despite its effectiveness, reward models (RMs) trained with BT model loss as one-way classifiers are prone to over-optimization, losing generalizability to unseen inputs. In thi…

Cited by 0SourcePDFScholar
2024

Stable Language Model Pre-training by Reducing Embedding Variability

EMNLP 2024main

Stable pre-training is essential for achieving better-performing language models. However, tracking pre-training stability is impractical due to high computational costs. We study Token Embedding Variability as a simple proxy to estimate pre-training stability. We theoretically and empirically demon…

Cited by 0SourcePDFScholar
2023

Disentangling Structure and Style: Political Bias Detection in News by Inducing Document Hierarchy

EMNLP 2023long findings

We address an important gap in detecting political bias in news articles. Previous works that perform document classification can be influenced by the writing style of each news outlet, leading to overfitting and limited generalizability. Our approach overcomes this limitation by considering both th…

Cited by 0SourceScholar
2021

Activation Sharing with Asymmetric Paths Solves Weight Transport Problem without Bidirectional Connection

NeurIPS 2021poster

One of the reasons why it is difficult for the brain to perform backpropagation (BP) is the weight transport problem, which argues forward and feedback neurons cannot share the same synaptic weights during learning in biological neural networks. Recently proposed algorithms address the weight transp…

Cited by 2SourcePDFScholar