← Search

Pegah Jandaghi

5 accepted papers

2025

A Systematic Analysis of Base Model Choice for Reward Modeling

EMNLP 2025

Reinforcement learning from human feedback (RLHF) and, at its core, reward modeling have become a crucial part of training powerful large language models (LLMs). One commonly overlooked factor in training high-quality reward models (RMs) is the effect of the base model, which is becoming more challe

Cited by 0SourcePDFScholar
2024

Faithful Persona-based Conversational Dataset Generation with Large Language Models

ACL 2024findings

High-quality conversational datasets are essential for developing AI models that can communicate with users.One way to foster deeper interactions between a chatbot and its user is through *personas*, aspects of the user’s character that provide insights into their personality, motivations, and behav…

2022

FETA: A Benchmark for Few-Sample Task Transfer in Open-Domain Dialogue

EMNLP 2022main

Task transfer, transferring knowledge contained in related tasks, holds the promise of reducing the quantity of labeled data required to fine-tune language models. Dialogue understanding encompasses many diverse tasks, yet task transfer has not been thoroughly studied in conversational AI. This work…

2022

Reflect, Not Reflex: Inference-Based Common Ground Improves Dialogue Response Quality

EMNLP 2022main

Human communication relies on common ground (CG), the mutual knowledge and beliefs shared by participants, to produce coherent and interesting conversations. In this paper, we demonstrate that current response generation (RG) models produce generic and dull responses in dialogues because they act re…

Cited by 30SourcePDFScholar
2021

Probing Commonsense Explanation in Dialogue Response Generation

EMNLP 2021finding

Humans use commonsense reasoning (CSR) implicitly to produce natural and coherent responses in conversations. Aiming to close the gap between current response generation (RG) models and human communication abilities, we want to understand why RG models respond as they do by probing RG model’s unders…

Cited by 19SourcePDFScholar