← Search

Shweta Singh

3 accepted papers

2025

IPO: Your Language Model is Secretly a Preference Classifier

ACL 2025long

Reinforcement learning from human feedback (RLHF) has emerged as the primary method for aligning large language models (LLMs) with human preferences. While it enables LLMs to achieve human-level alignment, it often incurs significant computational and financial costs due to its reliance on training…

2025

Open-DeBias: Toward Mitigating Open-Set Bias in Language Models

EMNLP 2025

Large Language Models (LLMs) have achieved remarkable success on question answering (QA) tasks, yet they often encode harmful biases that compromise fairness and trustworthiness. Most existing bias mitigation approaches are restricted to predefined categories, limiting their ability to address novel

Cited by 0SourcePDFScholar
2024

Benchmarking Object Detectors with COCO: A New Path Forward

ECCV 2024poster

"The Common Objects in Context (COCO) dataset has been instrumental in benchmarking object detectors over the past decade. Like every dataset, COCO contains subtle errors and imperfections stemming from its annotation procedure. With the advent of high-performing models, we ask whether these errors…