← Search

Andy Xu

3 accepted papers

2026

PLaID++: A Preference Aligned Language Model for Targeted Inorganic Materials Design

ICML 2026poster

Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising approach to improve correctness in LLMs, however, in many scientific problems, the objective is not necessarily to produce \textit{the} correct answer, but instead to produce a diverse array of candidates which satisfy …

Cited by 0SourceScholar