← Search

Kundan Krishna

8 accepted papers

2026

Learning to Reason for Hallucination Span Detection

ICLR 2026poster

Large language models (LLMs) often generate hallucinations---unsupported content that undermines reliability. While most prior works frame hallucination detection as a binary task, many real-world applications require identifying hallucinated spans, which is a multi-step decision making process. Thi…

Cited by 0SourceScholar
2025

Position: Towards Bidirectional Human-AI Alignment

NeurIPS 2025poster

Recent advances in general-purpose AI underscore the urgent need to align AI systems with human goals and values. Yet, the lack of a clear, shared understanding of what constitutes "alignment" limits meaningful progress and cross-disciplinary collaboration. In this position paper, we argue that the…

Cited by 0SourceScholar
2023

Downstream Datasets Make Surprisingly Good Pretraining Corpora

ACL 2023long

For most natural language processing tasks, the dominant practice is to finetune large pretrained transformer models (e.g., BERT) using smaller downstream datasets. Despite the success of this approach, it remains unclear to what extent these gainsare attributable to the massive background corpora e…

2023

Improving the Robustness of Summarization Models by Detecting and Removing Input Noise

EMNLP 2023long findings

The evaluation of abstractive summarization models typically uses test data that is identically distributed as training data. In real-world practice, documents to be summarized may contain input noise caused by text extraction artifacts or data pipeline bugs. The robustness of model performance unde…

Cited by 0SourceScholar
2023

Out-of-Distribution Detection and Selective Generation for Conditional Language Models

ICLR 2023top-25%

Machine learning algorithms typically assume independent and identically distributed samples in training and at test time (IID). Much work has shown that high-performing ML classifiers can degrade significantly and provide overly-confident, wrong classification predictions, particularly for out-of-…

Cited by 106SourcePDFScholar
2023

USB: A Unified Summarization Benchmark Across Tasks and Domains

EMNLP 2023long findings

While the NLP community has produced numerous summarization benchmarks, none provide the rich annotations required to simultaneously address many important problems related to control and reliability. We introduce a Wikipedia-derived benchmark, complemented by a rich set of crowd-sourced annotatio…

Cited by 0SourcecodeScholar
2021

Does Pretraining for Summarization Require Knowledge Transfer?

EMNLP 2021finding

Pretraining techniques leveraging enormous datasets have driven recent advances in text summarization. While folk explanations suggest that knowledge transfer accounts for pretraining’s benefits, little is known about why it works or what makes a pretraining task or dataset suitable. In this paper,…

Cited by 45SourcePDFScholar
2021

Generating SOAP Notes from Doctor-Patient Conversations Using Modular Summarization Techniques

ACL 2021long

Following each patient visit, physicians draft long semi-structured clinical summaries called SOAP notes. While invaluable to clinicians and researchers, creating digital SOAP notes is burdensome, contributing to physician burnout. In this paper, we introduce the first complete pipelines to leverage…

Cited by 132SourcePDFScholar