← Search

Pu-Jen Cheng

4 accepted papers

2025

Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification

ICLR 2025spotlight

Synthetic data augmentation via Large Language Models (LLMs) allows researchers to leverage additional training data, thus enhancing the performance of downstream tasks, especially when real-world data is scarce. However, the generated data can deviate from the real-world data, and this misalignment…

Cited by 1SourcePDFScholar
2024

Plug-in Language Model: Controlling Text Generation with a Simple Regression Model

NAACL 2024findings

Large-scale pre-trained language models have displayed unrivaled capacity in generating text that closely resembles human-written text. Nevertheless, generating texts adhering to specific conditions without fine-tuning or adding new parameters can be challenging. Contemporary approaches commonly rel…

2023

Hierarchical Programmatic Reinforcement Learning via Learning to Compose Programs

ICML 2023poster

Aiming to produce reinforcement learning (RL) policies that are human-interpretable and can generalize better to novel scenarios, Trivedi et al. (2021) present a method (LEAPS) that first learns a program embedding space to continuously parameterize diverse programs from a pre-generated program data…

Cited by 18SourcePDFScholar