← Search

Bishwaranjan Bhattacharjee

6 accepted papers

2026

GneissWeb: Preparing High Quality Data for LLMs at Scale

ICLR 2026poster

Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of downstream tasks. In this paper, we introduce **GneissWeb**, a large dataset of aro…

Cited by 0SourceScholar
2025

A Simple-Yet-Efficient Instruction Augmentation Method for Zero-Shot Sentiment Classification

COLING 2025main

Instruction tuning significantly enhances the performance of large language models in tasks such as sentiment classification. Previous studies have leveraged labeled instances from sentiment benchmark datasets to instruction-tune LLMs, improving zero-shot sentiment classification performance. In thi…

2025

Bias Analysis and Mitigation through Protected Attribute Detection and Regard Classification

EMNLP 2025

Large language models (LLMs) acquire general linguistic knowledge from massive-scale pretraining. However, pretraining data mainly comprised of web-crawled texts contain undesirable social biases which can be perpetuated or even amplified by LLMs. In this study, we propose an efficient yet effective

Cited by 0SourcePDFScholar
2024

INDUS: Effective and Efficient Language Models for Scientific Applications

EMNLP 2024industry

Large language models (LLMs) trained on general domain corpora showed remarkable results on natural language processing (NLP) tasks. However, previous research demonstrated LLMs trained using domain-focused corpora perform better on specialized tasks. Inspired by this insight, we developed INDUS, a…

Cited by 8SourcePDFScholar
2023

A Simple Yet Strong Domain-Agnostic De-bias Method for Zero-Shot Sentiment Classification

ACL 2023findings

Zero-shot prompt-based learning has made much progress in sentiment analysis, and considerable effort has been dedicated to designing high-performing prompt templates. However, two problems exist; First, large language models are often biased to their pre-training data, leading to poor performance i…

Cited by 7SourcePDFScholar
2021

NASTransfer: Analyzing Architecture Transferability in Large Scale Neural Architecture Search

AAAI 2021technical

Neural Architecture Search (NAS) is an open and challenging problem in machine learning. While NAS offers great promise, the prohibitive computational demand of most of the existing NAS methods makes it difficult to directly search the architectures on large-scale tasks. The typical way of conductin…

Cited by 13SourcePDFScholar