← Search

Yiyang Zhao

6 accepted papers

2026

GEM: Generative Entropy-Guided Preference Modeling for Few-Shot Alignment of LLMs

AAAI 2026technical

Alignment of large language models (LLMs) with human preferences typically relies on supervised reward models or external judges that demand abundant annotations. However, in fields that rely on professional knowledge, such as medicine and law, such large-scale preference labels are often unachievab

Cited by 0SourcePDFScholar
2026

GeoAlign: Geometric Rollout Curation for Robust LLM Reinforcement Learning

ICML 2026poster

Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under noisy or misspecified rewards. We identify a failure mode we call directional inconsistency: within a batch, a small set of high-reward rollouts induces represen…

Cited by 0SourceScholar
2024

CE-NAS: An End-to-End Carbon-Efficient Neural Architecture Search Framework

NeurIPS 2024poster

This work presents a novel approach to neural architecture search (NAS) that aims to increase carbon efficiency for the model design process. The proposed framework CE-NAS addresses the key challenge of high carbon cost associated with NAS by exploring the carbon emission variations of energy and en…

2022

Multi-objective Optimization by Learning Space Partition

ICLR 2022poster

In contrast to single-objective optimization (SOO), multi-objective optimization (MOO) requires an optimizer to find the Pareto frontier, a subset of feasible solutions that are not dominated by other feasible solutions. In this paper, we propose LaMOO, a novel multi-objective optimizer that learns…

Cited by 30SourcePDFScholar