← Search

Kokil Jaidka

13 accepted papers

2026

Always Refuse: Steering LLMs Against Jailbreaks with Contrastive Activations (Student Abstract)

AAAI 2026technical

Refusals must be resilient, not brittle.” Yet guarding refusals against adversarial phrasing and shifting user contexts remains difficult: large language models (LLMs) still yield to jailbreak prompts that evade safety filters and surface harmful content. We propose Refusal Activation Steering (RAS)

Cited by 0SourcePDFScholar
2026

Predicting Session Termination and Retention on X from Fine-Grained Interaction Logs (Student Abstract)

AAAI 2026technical

We study when users end a session on X using high-resolution interaction logs from 215 US participants collected over four weeks. Sessions are defined via data-driven inter-activity gaps, and each session is encoded by fine-grained activity counts and duration (versus a simple activity ratio baselin

Cited by 0SourcePDFScholar
2025

Beyond Context to Cognitive Appraisal: Emotion Reasoning as a Theory of Mind Benchmark for Large Language Models

ACL 2025finding

Datasets used for emotion recognition tasks typically contain overt cues that can be used in predicting the emotions expressed in a text. However, one challenge is that texts sometimes contain covert contextual cues that are rich in affective semantics, which warrant higher-order reasoning abilities…

Cited by 0SourcePDFScholar
2025

Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling (Student Abstract)

AAAI 2025technical

The prevalence of multi-modal content on social media complicates automated moderation strategies. This calls for an enhancement in multi-modal classification and a deeper understanding of understated meanings in images and memes. Although previous efforts have aimed at improving model performance t…

2025

Understanding Annotator Perception: Modeling Psychological Inference from First- and Third-Person Annotations (Student Abstract)

AAAI 2025technical

Large language models (LLMs) are trained on vast amounts of publicly available text. However, the current training frameworks take for granted that these annotations are accurate reflections of the authors’ true intents. This study questions that assumption by examining the gaps between writers’ act…

Cited by 0SourcePDFScholar
2024

Beyond Text: Leveraging Multi-Task Learning and Cognitive Appraisal Theory for Post-Purchase Intention Analysis

ACL 2024findings

Supervised machine-learning models for predicting user behavior offer a challenging classification problem with lower average prediction performance scores than other text classification tasks. This study evaluates multi-task learning frameworks grounded in Cognitive Appraisal Theory to predict user…

Cited by 1SourcePDFScholar
2024

Evaluating the Efficacy of Prompting Techniques for Debiasing Language Model Outputs (Student Abstract)

AAAI 2024technical

Achieving fairness in Large Language Models (LLMs) continues to pose a persistent challenge, as these models are prone to inheriting biases from their training data, which can subsequently impact their performance in various applications. There is a need to systematically explore whether structured…

Cited by 2SourcePDFScholar
2024

“Thinking” Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models

EMNLP 2024main

Existing debiasing techniques are typically training-based or require access to the model’s internals and output distributions, so they are inaccessible to end-users looking to adapt LLM outputs for their particular needs. In this study, we examine whether structured prompting techniques can offer o…

Cited by 11SourcePDFScholar
2023

Quantify the Political Bias in News Edits: Experiments with Few-Shot Learners (Student Abstract)

AAAI 2023technical

The rapid growth of information and communication technologies in recent years, and the different forms of digital connectivity, have profoundly affected how news is generated and consumed. Digital traces and computational methods offer new opportunities to model and track the provenance of news. Th…

Cited by 1SourcePDFScholar
2023

The PEACE-Reviews dataset: Modeling Cognitive Appraisals in Emotion Text Analysis

EMNLP 2023long findings

Cognitive appraisal plays a pivotal role in deciphering emotions. Recent studies have delved into its significance, yet the interplay between various forms of cognitive appraisal and specific emotions, such as joy and anger, remains an area of exploration in consumption contexts. Our research introd…

Cited by 0SourceScholar
2022

Offer a Different Perspective: Modeling the Belief Alignment of Arguments in Multi-party Debates

EMNLP 2022main

In contexts where debate and deliberation are the norm, the participants are regularly presented with new information that conflicts with their original beliefs. When required to update their beliefs (belief alignment), they may choose arguments that align with their worldview (confirmation bias). W…

2021

WikiTalkEdit: A Dataset for modeling Editors’ behaviors on Wikipedia

NAACL 2021long

This study introduces and analyzes WikiTalkEdit, a dataset of conversations and edit histories from Wikipedia, for research in online cooperation and conversation modeling. The dataset comprises dialog triplets from the Wikipedia Talk pages, and editing actions on the corresponding articles being di…