← Search

Saurav Sahay

3 accepted papers

2025

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

EMNLP 2025

Fine-tuning large language models (LLMs) for downstream tasks often leads to catastrophic forgetting, notably degrading the safety of originally aligned models. While some existing methods attempt to restore safety by incorporating additional safety data, the quality of such data typically falls sho

Cited by 0SourcePDFScholar
2021

Put Chatbot into Its Interlocutor’s Shoes: New Framework to Learn Chatbot Responding with Intention

NAACL 2021long

Most chatbot literature that focuses on improving the fluency and coherence of a chatbot, is dedicated to making chatbots more human-like. However, very little work delves into what really separates humans from chatbots – humans intrinsically understand the effect their responses have on the interlo…

Cited by 7SourcePDFScholar
2021

Refine and Imitate: Reducing Repetition and Inconsistency in Persuasion Dialogues via Reinforcement Learning and Human Demonstration

EMNLP 2021finding

Persuasion dialogue system reflects the machine’s ability to make strategic moves beyond verbal communication, and therefore differentiates itself from task-oriented or open-domain dialogues and has its own unique values. However, the repetition and inconsistency problems still persist in dialogue r…

Cited by 32SourcePDFScholar