← Search

Huiyu Bai

1 accepted papers

2026

GEM: Generative Entropy-Guided Preference Modeling for Few-Shot Alignment of LLMs

AAAI 2026technical

Alignment of large language models (LLMs) with human preferences typically relies on supervised reward models or external judges that demand abundant annotations. However, in fields that rely on professional knowledge, such as medicine and law, such large-scale preference labels are often unachievab

Cited by 0SourcePDFScholar