← Search

Jiajian Guo

2 accepted papers

2025

Advantage-Guided Distillation for Preference Alignment in Small Language Models

ICLR 2025spotlight

Alignment techniques enable Large Language Models (LLMs) to generate outputs that align with human preferences and play a crucial role in their effectiveness. However, their impact often diminishes when applied to Small Language Models (SLMs), likely due to the limited capacity of these models. Inst…

2025

ReAlign: Structured Revision for Small Language Model Alignment

EMNLP 2025

Aligning small language models with human preferences is challenging, as weak policies struggle to generate informative on-policy samples and suffer from unstable gradients when trained on off-policy signals from stronger models. In this work, we propose ReAlign, a training framework that combines t