← Search

Yi-Chern Tan

1 accepted papers

2025

Reverse Engineering Human Preferences with Reinforcement Learning

NeurIPS 2025spotlight

The capabilities of Large Language Models (LLMs) are routinely evaluated by other LLMs trained to predict human preferences. This framework—known as *LLM-as-a-judge*—is highly scalable and relatively low cost. However, it is also vulnerable to malicious exploitation, as LLM responses can be tuned to…

Cited by 0SourceScholar