← Search

Jad Kabbara

16 accepted papers

2026

LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users

AAAI 2026technical

While state-of-the-art large language models (LLMs) have shown impressive performance on many tasks, systematically evaluating undesirable behaviors of these models remains critical. In this work, we investigate how the quality of LLM responses changes in terms of information accuracy, truthfulness,

Cited by 0SourcePDFScholar
2025

Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts

COLING 2025main

We focus on enhancing comprehension in small-group recorded conversations, which serve as a medium to bring people together and provide a space for sharing personal stories and experiences on crucial social matters. One way to parse and convey information from these conversations is by sharing highl…

2025

Bridging the Data Provenance Gap Across Text, Speech, and Video

ICLR 2025poster

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and first-of-its-kind longitudinal audit across modalities --- pop…

Cited by 1SourcePDFScholar
2025

Computational Analysis of Conversation Dynamics through Participant Responsivity

EMNLP 2025

Growing literature explores toxicity and polarization in discourse, with comparatively less work on characterizing what makes dialogue prosocial and constructive. We explore conversational discourse and investigate a method for characterizing its quality built upon the notion of “responsivity”—wheth

Cited by 0SourcePDFScholar
2025

Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks

ACL 2025finding

LLM use in annotation is becoming widespread, and given LLMs’ overall promising performance and speed, putting humans in the loop to simply “review” LLM annotations can be tempting. In subjective tasks with multiple plausible answers, this can impact both evaluation of LLM performance, and analysis…

Cited by 0SourcePDFScholar
2024

Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models

ACL 2024long

As the use of Large Language Models (LLMs) becomes more widespread, understanding their self-evaluation of confidence in generated responses becomes increasingly important as it is integral to the reliability of the output of these models. We introduce the concept of Confidence-Probability Alignment…

2024

Consent in Crisis: The Rapid Decline of the AI Data Commons

NeurIPS 2024poster

General-purpose artificial intelligence (AI) systems are built on massive swathes of public web data, assembled into corpora such as C4, RefinedWeb, and Dolma. To our knowledge, we conduct the first, large-scale, longitudinal audit of the consent protocols for the web domains underlying AI training…

Cited by 36SourceScholar
2024

Leveraging Large Language Models for Learning Complex Legal Concepts through Storytelling

ACL 2024long

Making legal knowledge accessible to non-experts is crucial for enhancing general legal literacy and encouraging civic participation in democracy. However, legal documents are often challenging to understand for people without legal backgrounds. In this paper, we present a novel application of large…

2024

On the Relationship between Truth and Political Bias in Language Models

EMNLP 2024main

Language model alignment research often attempts to ensure that models are not only helpful and harmless, but also truthful and unbiased. However, optimizing these objectives simultaneously can obscure how improving one aspect might impact the others. In this work, we focus on analyzing the relation…

2024

PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits

NAACL 2024findings

Despite the many use cases for large language models (LLMs) in creating personalized chatbots, there has been limited research on evaluating the extent to which the behaviors of personalized LLMs accurately and consistently reflect specific personality traits. We consider studying the behavior of LL…

2024

Position: Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?

ICML 2024spotlight

New capabilities in foundation models are owed in large part to massive, widely-sourced, and under-documented training data collections. Existing practices in data collection have led to challenges in tracing authenticity, verifying consent, preserving privacy, addressing representation and bias, re…

Cited by 3SourcePDFScholar
2023

Debiasing should be Good and Bad: Measuring the Consistency of Debiasing Techniques in Language Models

ACL 2023findings

Debiasing methods that seek to mitigate the tendency of Language Models (LMs) to occasionally output toxic or inappropriate text have recently gained traction. In this paper, we propose a standardized protocol which distinguishes methods that yield not only desirable results, but are also consistent…

2023

Investigating the Effect of Pre-finetuning BERT Models on NLI Involving Presuppositions

EMNLP 2023long findings

We explore the connection between presupposition, discourse and sarcasm and propose to leverage that connection in a transfer learning scenario with the goal of improving the performance of NLI models on cases involving presupposition. We exploit advances in training transformer-based models that sh…

Cited by 0SourceScholar
2022

Investigating the Performance of Transformer-Based NLI Models on Presuppositional Inferences

COLING 2022main

Presuppositions are assumptions that are taken for granted by an utterance, and identifying them is key to a pragmatic interpretation of language. In this paper, we investigate the capabilities of transformer models to perform NLI on cases involving presupposition. First, we present simple heuristic…

Cited by 7SourcePDFScholar