← Search

Simerjot Kaur

7 accepted papers

2025

A Variational Approach for Mitigating Entity Bias in Relation Extraction

ACL 2025short

Mitigating entity bias is a critical challenge in Relation Extraction (RE), where models often rely excessively on entities, resulting in poor generalization. This paper presents a novel approach to address this issue by adapting a Variational Information Bottleneck (VIB) framework. Our method compr…

Cited by 0SourcePDFScholar
2025

Calibrating LLM Confidence by Probing Perturbed Representation Stability

EMNLP 2025

Miscalibration in Large Language Models (LLMs) undermines their reliability, highlighting the need for accurate confidence estimation. We introduce CCPS (Calibrating LLM Confidence by Probing Perturbed Representation Stability), a novel method analyzing internal representational stability in LLMs. C

2025

Conservative Bias in Large Language Models: Measuring Relation Predictions

ACL 2025finding

Large language models (LLMs) exhibit pronounced conservative bias in relation extraction tasks, frequently defaulting to no_relation label when an appropriate option is unavailable. While this behavior helps prevent incorrect relation assignments, our analysis reveals that it also leads to significa…

Cited by 0SourcePDFScholar
2025

FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking

NAACL 2025findings

We introduce FinNLI, a benchmark dataset for Financial Natural Language Inference (FinNLI) across diverse financial texts like SEC Filings, Annual Reports, and Earnings Call transcripts. Our dataset framework ensures diverse premise-hypothesis pairs while minimizing spurious correlations. FinNLI com…

Cited by 1SourcePDFScholar
2025

Translating Domain-Specific Terminology in Typologically-Diverse Languages: A Study in Tax and Financial Education

EMNLP 2025

Domain-specific multilingual terminology is essential for accurate machine translation (MT) and cross-lingual NLP applications. We present a gold-standard terminology resource for the tax and financial education domains, built from curated governmental publications and covering seven typologically d

Cited by 0SourcePDFScholar
2024

DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding

ACL 2024long

Enterprise documents such as forms, receipts, reports, and other such records, often carry rich semantics at the intersection of textual and spatial modalities. The visual cues offered by their complex layouts play a crucial role in comprehending these documents effectively. In this paper, we presen…

2024

Large Language Models as Financial Data Annotators: A Study on Effectiveness and Efficiency

COLING 2024main

Collecting labeled datasets in finance is challenging due to scarcity of domain experts and higher cost of employing them. While Large Language Models (LLMs) have demonstrated remarkable performance in data annotation tasks on general domain datasets, their effectiveness on domain specific datasets…

Cited by 20SourcePDFScholar