← Search

Tanmoy Chakraborty

59 accepted papers

2026

Beyond Markovian Drifts: Action-Biased Geometric Walks with Memory for Personalized Summarization

ICLR 2026poster

Document summarization helps readers focus on the "content-of-interest", a *subjective* and *time-variant* quantity. Capturing this *dynamic subjectivity* requires modeling how user preferences evolve over time, thereby demanding *personalized summarization*. Recent news recommendation and summariza…

Cited by 0SourceScholar
2026

Linguistic Properties and Model Scale in Brain Encoding: From Small to Compressed Language Models

ICML 2026spotlight

Recent work has shown that scaling large language models (LLMs) improves their alignment with human brain activity, yet it remains unclear what drives these gains or which representational properties are responsible. Although larger models often yield better task performance and brain alignment, the…

Cited by 0SourceScholar
2026

Polaris: Coupled Orbital Polar Embeddings for Hierarchical Concept Learning

ICML 2026poster

Real-world knowledge is often organized as hierarchies such as product taxonomies, medical ontologies, and label trees, yet learning hierarchical representations is challenging due to asymmetric structure and noisy semantics. We introduce Polaris, a polar hyperspherical embedding framework that sepa…

Cited by 0SourceScholar
2026

QuanTaxo: A Quantum Approach to Self-Supervised Taxonomy Expansion

AAAI 2026technical

A taxonomy is a hierarchical graph containing knowledge to provide valuable insights for various web applications. However, the manual construction of taxonomies requires significant human effort. As web content continues to expand at an unprecedented pace, existing taxonomies risk becoming outdated

Cited by 0SourcePDFScholar
2025

CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs

NAACL 2025long

Counterspeech has emerged as a popular and effective strategy for combating online hate speech, sparking growing research interest in automating its generation using language models. However, the field still lacks standardised evaluation protocols and reliable automated evaluation metrics that align…

2025

Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning

ACL 2025long

Counterspeech has proven to be a powerful tool to combat hate speech online. Previous studies have focused on generating counterspeech conditioned only on specific intents (single attributed). However, a holistic approach considering multiple attributes simultaneously can yield more nuanced and effe…

2025

Do You Know About My Nation? Investigating Multilingual Language Models’ Cultural Literacy Through Factual Knowledge

EMNLP 2025

Most multilingual question-answering benchmarks, while covering a diverse pool of languages, do not factor in regional diversity in the information they capture and tend to be Western-centric. This introduces a significant gap in fairly evaluating multilingual models’ comprehension of factual inform

2025

Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages

AAAI 2025technical

The unwavering disparity in labeled resources between resource-rich languages and those considered low-resource remains a significant impediment for Large Language Models (LLMs). Recent strides in cross-lingual in-context learning (X-ICL), mainly through semantically aligned examples retrieved from…

2025

Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models

NAACL 2025long

While recent large language models (LLMs) demonstrate remarkable abilities in responding to queries in diverse languages, their ability to handle long multilingual contexts is unexplored. As such, a systematic evaluation of the long-context capabilities of LLMs in multilingual settings is crucial, s…

2025

On the Generalization vs Fidelity Paradox in Knowledge Distillation

ACL 2025finding

Knowledge distillation (KD) is a key technique for compressing large language models into smaller ones while preserving performance. Despite the recent traction of KD research, its effectiveness for smaller language models (LMs) and the mechanisms driving knowledge transfer remain underexplored. In…

2025

Parallel Communities Across the Surface Web and the Dark Web

EMNLP 2025

Humans have an inherent need for community belongingness. This paper investigates this fundamental social motivation by compiling a large collection of parallel datasets comprising over 7 million posts and comments from Reddit and 200,000 posts and comments from Dread, a dark web discussion forum, c

2025

Rank, Chunk and Expand: Lineage-Oriented Reasoning for Taxonomy Expansion

ACL 2025finding

Taxonomies are hierarchical knowledge graphs crucial for recommendation systems, and web applications. As data grows, expanding taxonomies is essential, but existing methods face key challenges: (1) discriminative models struggle with representation limits and generalization, while (2) generative me…

2025

SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection

EMNLP 2025

Large Language Models (LLMs) with safe-alignment training are powerful instruments with robust language comprehension capability. Typically LLMs undergo careful alignment training involving human feedback to ensure the acceptance of safe inputs while rejection of harmful or unsafe ones. However, the

2025

SAHAY: Multimodal, Privacy-Preserving AI for Suicide Risk Detection and Intervention in India

IJCAI 2025

Suicide accounts for one of the leading causes of death in India, with over 164,033 deaths reported in 2021. Despite increased awareness, the gap between the need for consistent treatment and actual accessibility remains a challenge due to limited mental health infrastructure, the stigma surrounding

Cited by 2SourcePDFScholar
2025

Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?

EMNLP 2025

The increasing acceptance of large language models (LLMs) as an alternative to knowledge sources marks a significant paradigm shift across various domains, including time-sensitive fields such as law, healthcare, and finance. To fulfill this expanded role, LLMs must not only be factually accurate bu

2025

The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection

EMNLP 2025

Misinformation remains one of the most significant issues in the digital age. While automated fact-checking has emerged as a viable solution, most current systems are limited to evaluating factual accuracy. However, the detrimental effect of misinformation transcends simple falsehoods; it takes adva

Cited by 0SourcePDFScholar
2025

Value-Guided KV Compression for LLMs via Approximated CUR Decomposition

NeurIPS 2025poster

Key-value (KV) cache compression has emerged as a critical technique for reducing the memory and latency overhead of autoregressive language models during inference. Prior approaches predominantly rely on query-key attention scores to rank and evict cached tokens, assuming that attention intensity c…

Cited by 0SourceScholar
2025

You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning

ICLR 2025poster

The ever-increasing size of large language models (LLMs) presents significant challenges for deployment due to their heavy computational and memory requirements. Current model pruning techniques attempt to alleviate these issues by relying heavily on external calibration datasets to determine which…

Cited by 0SourcePDFScholar
2024

A Good Learner can Teach Better: Teacher-Student Collaborative Knowledge Distillation

ICLR 2024poster

Knowledge distillation (KD) is a technique used to transfer knowledge from a larger ''teacher'' model into a smaller ''student'' model. Recent advancements in meta-learning-based knowledge distillation (MetaKD) emphasize that the fine-tuning of teacher models should be aware of the student's need to…

Cited by 5SourcePDFScholar
2024

Adding SPICE to Life: Speaker Profiling in Multiparty Conversations

COLING 2024main

In the realm of conversational dynamics, individual idiosyncrasies challenge the suitability of a one-size-fits-all approach for dialogue agent responses. Prior studies often assumed the speaker’s persona’s immediate availability, a premise not universally applicable. To address this gap, we explore…

Cited by 1SourcePDFScholar
2024

Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators

EMNLP 2024main

Large Language Models (LLMs) and AI assistants driven by these models are experiencing exponential growth in usage among both expert and amateur users. In this work, we focus on evaluating the reliability of current LLMs as science communicators. Unlike existing benchmarks, our approach emphasizes a…

2024

Frugal LMs Trained to Invoke Symbolic Solvers Achieve Parameter-Efficient Arithmetic Reasoning

AAAI 2024technical

Large Language Models (LLM) exhibit zero-shot mathematical reasoning capacity as a behavior emergent with scale, commonly manifesting as chain-of-thoughts (CoT) reasoning. However, multiple empirical findings suggest that this prowess is exclusive to LLMs that have exorbitant sizes (beyond 50 billio…

2024

Hate Personified: Investigating the role of LLMs in content moderation

EMNLP 2024main

For subjective tasks such as hate detection, where people perceive hate differently, the Large Language Model’s (LLM) ability to represent diverse groups is unclear. By including additional context in prompts, we comprehensively analyze LLM’s sensitivity to geographical priming, persona attributes,…

2024

Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF

NAACL 2024long

Counterspeech, defined as a response to mitigate online hate speech, is increasingly used as a non-censorial solution. The effectiveness of addressing hate speech involves dispelling the stereotypes, prejudices, and biases often subtly implied in brief, single-sentence statements or abuses. These ex…

2024

Knowledge Planning in Large Language Models for Domain-Aligned Counseling Summarization

EMNLP 2024main

In mental health counseling, condensing dialogues into concise and relevant summaries (aka counseling notes) holds pivotal significance. Large Language Models (LLMs) exhibit remarkable capabilities in various generative tasks; however, their adaptation to domain-specific intricacies remains challeng…

2024

LM2: A Simple Society of Language Models Solves Complex Reasoning

EMNLP 2024main

Despite demonstrating emergent reasoning abilities, Large Language Models (LLMS) often lose track of complex, multi-step reasoning. Existing studies show that providing guidance via decomposing the original question into multiple subproblems elicits more robustness in LLM reasoning – a decomposer ge…

Cited by 1SourcePDFScholar
2024

Language Models can Exploit Cross-Task In-context Learning for Data-Scarce Novel Tasks

ACL 2024long

Large Language Models (LLMs) have transformed NLP with their remarkable In-context Learning (ICL) capabilities. Automated assistants based on LLMs are gaining popularity; however, adapting them to novel tasks is still challenging. While colossal models excel in zero-shot performance, their computati…

2024

MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing

ACL 2024findings

Memes have evolved as a prevalent medium for diverse communication, ranging from humour to propaganda. With the rising popularity of image-focused content, there is a growing need to explore its potential harm from different aspects. Previous studies have analyzed memes in closed settings - detectin…

2024

POSIX: A Prompt Sensitivity Index For Large Language Models

EMNLP 2024finding

Despite their remarkable capabilities, Large Language Models (LLMs) are found to be surprisingly sensitive to minor variations in prompts, often generating significantly divergent outputs in response to minor variations in the prompts, such as spelling errors, alteration of wording or the prompt tem…

2024

Recent Advances in Online Hate Speech Moderation: Multimodality and the Role of Large Models

EMNLP 2024finding

Moderating hate speech (HS) in the evolving online landscape is a complex challenge, compounded by the multimodal nature of digital content. This survey examines recent advancements in HS moderation, focusing on the burgeoning role of large language models (LLMs) and large multimodal models (LMMs) i…

Cited by 1SourcePDFScholar
2024

SUKHSANDESH: An Avatar Therapeutic Question Answering Platform for Sexual Education in Rural India

IJCAI 2024poster

Sexual education aims to foster a healthy lifestyle in terms of emotional, mental and social well-being. In countries like India, where adolescents form the largest demographic group, they face significant vulnerabilities concerning sexual health. Unfortunately, sexual education is often stigmatize…

Cited by 0SourcePDFScholar
2024

Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation

COLING 2024main

The widespread online communication in a modern multilingual world has provided opportunities to blend more than one language (aka code-mixed language) in a single utterance. This has resulted a formidable challenge for the computational models due to the scarcity of annotated data and presence of n…

2024

Temporally Consistent Factuality Probing for Large Language Models

EMNLP 2024main

The prolific use of Large Language Models (LLMs) as an alternate knowledge base requires them to be factually consistent, necessitating both correctness and consistency traits for paraphrased queries. Recently, significant attempts have been made to benchmark datasets and metrics to evaluate LLMs fo…

2024

Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech

ACL 2024findings

Employing language models to generate explanations for an incoming implicit hate post is an active area of research. The explanation is intended to make explicit the underlying stereotype and aid content moderators. The training often combines top-k relevant knowledge graph (KG) tuples to provide wo…

2023

$\textit{From Chaos to Clarity}$: Claim Normalization to Empower Fact-Checking

EMNLP 2023long findings

With the rise of social media, users are exposed to many misleading claims. However, the pervasive noise inherent in these posts presents a challenge in identifying precise and prominent claims that require verification. Extracting the important claims from such posts is arduous and time-consuming,…

Cited by 0SourceScholar
2023

Accuracy is not enough: Evaluating Personalization in Summarizers

EMNLP 2023long findings

Text summarization models are evaluated in terms of their accuracy and quality using various measures such as ROUGE, BLEU, METEOR, BERTScore, PYRAMID, readability, and several other recently proposed ones. The central objective of all accuracy measures is to evaluate the model's ability to capture $…

Cited by 0SourceScholar
2023

Counterspeeches up my sleeve! Intent Distribution Learning and Persistent Fusion for Intent-Conditioned Counterspeech Generation

ACL 2023long

Counterspeech has been demonstrated to be an efficacious approach for combating hate speech. While various conventional and controlled approaches have been studied in recent years to generate counterspeech, a counterspeech with a certain intent may not be sufficient in every scenario. Due to the com…

2023

Explaining (Sarcastic) Utterances to Enhance Affect Understanding in Multimodal Dialogues

AAAI 2023technical

Conversations emerge as the primary media for exchanging ideas and conceptions. From the listener’s perspective, identifying various affective qualities, such as sarcasm, humour, and emotions, is paramount for comprehending the true connotation of the emitted utterance. However, one of the major hur…

2023

From Multilingual Complexity to Emotional Clarity: Leveraging Commonsense to Unveil Emotions in Code-Mixed Dialogues

EMNLP 2023long main

Understanding emotions during conversation is a fundamental aspect of human communication, driving NLP research for Emotion Recognition in Conversation (ERC). While considerable research has focused on discerning emotions of individual speakers in monolingual dialogues, understanding the emotional d…

Cited by 0SourcecodeScholar
2023

Learning and Reasoning Multifaceted and Longitudinal Data for Poverty Estimates and Livelihood Capabilities of Lagged Regions in Rural India

IJCAI 2023poster

Poverty is a multifaceted phenomenon linked to the lack of capabilities of households to earn a sustainable livelihood, increasingly being assessed using multidimensional indicators. Its spatial pattern depends on social, economic, political, and regional variables. Artificial intelligence has shown…

Cited by 2SourcePDFScholar
2023

MEMEX: Detecting Explanatory Evidence for Memes via Knowledge-Enriched Contextualization

ACL 2023long

Memes are a powerful tool for communication over social media. Their affinity for evolving across politics, history, and sociocultural phenomena renders them an ideal vehicle for communication. To comprehend the subtle message conveyed within a meme, one must understand the relevant background that…

2023

Manifold-Preserving Transformers are Effective for Short-Long Range Encoding

EMNLP 2023long findings

Multi-head self-attention-based Transformers have shown promise in different learning tasks. Albeit these models exhibit significant improvement in understanding short-term and long-term contexts from sequences, encoders of Transformers and their variants fail to preserve layer-wise contextual infor…

Cited by 0SourcecodeScholar
2023

Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment

ACL 2023long

In-context learning (ICL) unfolds as large language models become capable of inferring test labels conditioned on a few labeled samples without any gradient update. ICL-enabled large language models provide a promising step forward toward bypassing recurrent annotation costs in a low-resource settin…

2023

Promoting Topic Coherence and Inter-Document Consorts in Multi-Document Summarization via Simplicial Complex and Sheaf Graph

EMNLP 2023long main

Multi-document Summarization (MDS) characterizes compressing information from multiple source documents to its succinct summary. An ideal summary should encompass all topics and accurately model cross-document relations expounded upon in the source documents. However, existing systems either impose…

Cited by 0SourceScholar
2023

Small Language Models Fine-tuned to Coordinate Larger Language Models improve Complex Reasoning

EMNLP 2023long main

Large Language Models (LLMs) prompted to generate chain-of-thought (CoT) exhibit impressive reasoning capabilities. Recent attempts at prompt decomposition toward solving complex, multi-step reasoning problems depend on the ability of the LLM to simultaneously decompose and solve the problem. A sign…

Cited by 0SourcecodeScholar
2023

What Do You MEME? Generating Explanations for Visual Semantic Role Labelling in Memes

AAAI 2023technical

Memes are powerful means for effective communication on social media. Their effortless amalgamation of viral visuals and compelling messages can have far-reaching implications with proper marketing. Previous research on memes has primarily focused on characterizing their affective spectrum and detec…

2022

A Survey on Multimodal Disinformation Detection

COLING 2022main

Recent years have witnessed the proliferation of offensive content online such as fake news, propaganda, misinformation, and disinformation. While initially this was mostly about textual content, over time images and videos gained popularity, as they are much easier to consume, attract more attentio…

Cited by 156SourcePDFScholar
2022

Can Unsupervised Knowledge Transfer from Social Discussions Help Argument Mining?

ACL 2022long

Identifying argument components from unstructured texts and predicting the relationships expressed among them are two primary steps of argument mining. The intrinsic complexity of these tasks demands powerful learning models. While pretrained Transformer-based Language Models (LM) have been shown to…

2022

DISARM: Detecting the Victims Targeted by Harmful Memes

NAACL 2022findings

Internet memes have emerged as an increasingly popular means of communication on the web. Although memes are typically intended to elicit humour, they have been increasingly used to spread hatred, trolling, and cyberbullying, as well as to target specific individuals, communities, or society on poli…

2022

Detecting and Understanding Harmful Memes: A Survey

IJCAI 2022poster

The automatic identification of harmful content online is of major concern for social media platforms, policymakers, and society. Researchers have studied textual, visual, and audio content, but typically in isolation. Yet, harmful content often combines multiple modalities, as in the case of memes.…

2022

Empowering the Fact-checkers! Automatic Identification of Claim Spans on Twitter

EMNLP 2022main

The widespread diffusion of medical and political claims in the wake of COVID-19 has led to a voluminous rise in misinformation and fake news. The current vogue is to employ manual fact-checkers to efficiently classify and verify such data to combat this avalanche of claim-ridden misinformation. How…

2022

Nice Perfume. How Long Did You Marinate in It? Multimodal Sarcasm Explanation

AAAI 2022technical

Sarcasm is a pervading linguistic phenomenon and highly challenging to explain due to its subjectivity, lack of context and deeply-felt opinion. In the multimodal setup, sarcasm is conveyed through the incongruity between the text and visual entities. Although recent approaches deal with sarcasm as…

2022

Public Wisdom Matters! Discourse-Aware Hyperbolic Fourier Co-Attention for Social Text Classification

NeurIPS 2022accept

Social media has become the fulcrum of all forms of communication. Classifying social texts such as fake news, rumour, sarcasm, etc. has gained significant attention. The surface-level signals expressed by a social-text itself may not be adequate for such tasks; therefore, recent methods attempted t…

2022

When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues

ACL 2022long

Indirect speech such as sarcasm achieves a constellation of discourse goals in human communication. While the indirectness of figurative language warrants speakers to achieve certain pragmatic goals, it is challenging for AI agents to comprehend such idiosyncrasies of human communication. Though sar…

2021

MOMENTA: A Multimodal Framework for Detecting Harmful Memes and Their Targets

EMNLP 2021finding

Internet memes have become powerful means to transmit political, psychological, and socio-cultural ideas. Although memes are typically humorous, recent days have witnessed an escalation of harmful memes used for trolling, cyberbullying, and abuse. Detecting such memes is challenging as they can be h…

2021

Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical Systems

NeurIPS 2021spotlight

The Transformer and its variants have been proven to be efficient sequence learners in many different domains. Despite their staggering success, a critical issue has been the enormous number of parameters that must be trained (ranging from $10^7$ to $10^{11}$) along with the quadratic complexity of…

2020

Neural Abstractive Summarization with Structural Attention

IJCAI 2020poster

Attentional, RNN-based encoder-decoder architectures have obtained impressive performance on abstractive summarization of news articles. However, these methods fail to account for long term dependencies within the sentences of a document. This problem is exacerbated in multi-document summarization t…