← Search

Zihuiwen Ye

2 accepted papers

2025

Improving Reward Models with Synthetic Critiques

NAACL 2025findings

Reward models (RMs) play a critical role in aligning language models through the process of reinforcement learning from human feedback. RMs are trained to predict a score reflecting human preference, which requires significant time and cost for human annotation. Additionally, RMs tend to quickly ove…

2022

Augmenting Multi-Turn Text-to-SQL Datasets with Self-Play

EMNLP 2022finding

The task of context-dependent text-to-SQL aims to convert multi-turn user utterances to formal SQL queries. This is a challenging task due to both the scarcity of training data from which to learn complex contextual dependencies and to generalize to unseen databases. In this paper we explore augment…