← Search

Sudipta Kar

4 accepted papers

2025

Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models

AAAI 2025technical

Fine-tuning large language models (LLMs) with a collection of large and diverse instructions has improved the model’s generalization to different tasks, even for unseen tasks. However, most existing instruction datasets include only single instructions, and they struggle to follow complex instructio…

2023

MultiCoNER v2: a Large Multilingual dataset for Fine-grained and Noisy Named Entity Recognition

EMNLP 2023short findings

We present MULTICONER V2, a dataset for fine-grained Named Entity Recognition covering 33 entity classes across 12 languages, in both monolingual and multilingual settings. This dataset aims to tackle the following practical challenges in NER: (i) effective handling of fine-grained classes that incl…

Cited by 0SourceScholar
2022

MultiCoNER: A Large-scale Multilingual Dataset for Complex Named Entity Recognition

COLING 2022main

We present AnonData, a large multilingual dataset for Named Entity Recognition that covers 3 domains (Wiki sentences, questions, and search queries) across 11 languages, as well as multilingual and code-mixing subsets. This dataset is designed to represent contemporary challenges in NER, including l…

Cited by 114SourcePDFScholar
2021

SentNoB: A Dataset for Analysing Sentiment on Noisy Bangla Texts

EMNLP 2021finding

In this paper, we propose an annotated sentiment analysis dataset made of informally written Bangla texts. This dataset comprises public comments on news and videos collected from social media covering 13 different domains, including politics, education, and agriculture. These comments are labeled w…