← Search

Shikib Mehri

5 accepted papers

2025

LMUNIT: Fine-grained Evaluation with Natural Language Unit Tests

EMNLP 2025

As language models become integral to critical workflows, assessing their behavior remains a fundamental challenge – human evaluation is costly and noisy, while automated metrics provide only coarse, difficult-to-interpret signals. We introduce natural language unit tests , a paradigm that decompose

2023

CESAR: Automatic Induction of Compositional Instructions for Multi-turn Dialogs

EMNLP 2023long main

Instruction-based multitasking has played a critical role in the success of large language models (LLMs) in multi-turn dialog applications. While publicly available LLMs have shown promising performance, when exposed to complex instructions with multiple constraints, they lag against state-of-the-ar…

Cited by 0SourceScholar
2022

InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning

EMNLP 2022main

Instruction tuning is an emergent paradigm in NLP wherein natural language instructions are leveraged with language models to induce zero-shot performance on unseen tasks. Dialogue is an especially interesting area in which to explore instruction tuning because dialogue systems perform multiple kind…

2018

Middle-Out Decoding

NeurIPS 2018poster

Despite being virtually ubiquitous, sequence-to-sequence models are challenged by their lack of diversity and inability to be externally controlled. In this paper, we speculate that a fundamental shortcoming of sequence generation models is that the decoding is done strictly from left-to-right, mean…

Cited by 24SourcePDFScholar