← Search

Chenyan Jia

3 accepted papers

2024

Training Socially Aligned Language Models on Simulated Social Interactions

ICLR 2024poster

The goal of social alignment for AI systems is to make sure these models can conduct themselves appropriately following social values. Unlike humans who establish a consensus on value judgments through social interaction, current language models (LMs) are trained to rigidly recite the corpus in soci…

2022

Non-Parallel Text Style Transfer with Self-Parallel Supervision

ICLR 2022poster

The performance of existing text style transfer models is severely limited by the non-parallel datasets on which the models are trained. In non-parallel datasets, no direct mapping exists between sentences of the source and target style; the style transfer models thus only receive weak supervision o…

2022

Second Thoughts are Best: Learning to Re-Align With Human Values from Text Edits

NeurIPS 2022accept

We present Second Thoughts, a new learning paradigm that enables language models (LMs) to re-align with human values. By modeling the chain-of-edits between value-unaligned and value-aligned text, with LM fine-tuning and additional refinement through reinforcement learning, Second Thoughts not only…

Cited by 38SourcePDFScholar