← Search

Hal Daumé III

44 accepted papers

2025

Can Hallucination Correction Improve Video-Language Alignment?

ACL 2025finding

Large Vision-Language Models often generate hallucinated content that is not grounded in its visual inputs. While prior work focuses on mitigating hallucinations, we instead explore leveraging hallucination correction as a training objective to improve video-language alignment. We introduce HACA, a…

Cited by 0SourcePDFScholar
2025

Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation Learning

NeurIPS 2025spotlight

Language model activations entangle concepts that mediate their behavior, making it difficult to interpret these factors, which has implications for generalizability and robustness. We introduce an approach for disentangling these concepts without supervision. Existing methods for concept discovery…

Cited by 0SourceScholar
2025

Language Models Predict Empathy Gaps Between Social In-groups and Out-groups

NAACL 2025long

Studies of human psychology have demonstrated that people are more motivated to extend empathy to in-group members than out-group members (Cikara et al., 2011). In this study, we investigate how this aspect of intergroup relations in humans is replicated by LLMs in an emotion intensity prediction ta…

2025

My LLM might Mimic AAE - But When Should It?

NAACL 2025long

We examine the representation of African American English (AAE) in large language models (LLMs), exploring (a) the perceptions Black Americans have of how effective these technologies are at producing authentic AAE, and (b) in what contexts Black Americans find this desirable. Through both a survey…

2025

Natural Language Inference Improves Compositionality in Vision-Language Models

ICLR 2025poster

Compositional reasoning in Vision-Language Models (VLMs) remains challenging as these models often struggle to relate objects, attributes, and spatial relationships. Recent methods aim to address these limitations by relying on the semantics of the textual description, using Large Language Models (L…

Cited by 3SourcePDFScholar
2025

TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

ICLR 2025poster

Although large vision-language-action (VLA) models pretrained on extensive robot datasets offer promising generalist policies for robotic learning, they still struggle with spatial-temporal dynamics in interactive robotics, making them less effective in handling complex tasks, such as manipulation.…

Cited by 8SourcePDFScholar
2024

ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles

EMNLP 2024main

Deaf and hard-of-hearing (DHH) students face significant barriers in accessing science, technology, engineering, and mathematics (STEM) education, notably due to the scarcity of STEM resources in signed languages. To help address this, we introduce ASL STEM Wiki: a parallel corpus of 254 Wikipedia a…

Cited by 1SourcePDFScholar
2024

Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA

EMNLP 2024main

Recent advancements of large language models (LLMs)have led to claims of AI surpassing humansin natural language processing NLP tasks such as textual understanding and reasoning.%This work investigates these assertions by introducingCAIMIRA, a novel framework rooted in item response theory IRTthat e…

Cited by 2SourcePDFScholar
2024

DrM: Mastering Visual Reinforcement Learning through Dormant Ratio Minimization

ICLR 2024spotlight

Visual reinforcement learning (RL) has shown promise in continuous control tasks. Despite its progress, current algorithms are still unsatisfactory in virtually every aspect of the performance such as sample efficiency, asymptotic performance, and their robustness to the choice of random seeds. In t…

2024

HateCOT: An Explanation-Enhanced Dataset for Generalizable Offensive Speech Detection via Large Language Models

EMNLP 2024finding

The widespread use of social media necessitates reliable and efficient detection of offensive content to mitigate harmful effects. Although sophisticated models perform well on individual datasets, they often fail to generalize due to varying definitions and labeling of “offensive content.” In this…

2024

Large Language Models Help Humans Verify Truthfulness – Except When They Are Convincingly Wrong

NAACL 2024long

Large Language Models (LLMs) are increasingly used for accessing information on the web. Their truthfulness and factuality are thus of great interest. To help users make the right decisions about the information they get, LLMs should not only provide information but also help users fact-check it. We…

Cited by 36SourcePDFScholar
2024

PRISE: LLM-Style Sequence Compression for Learning Temporal Action Abstractions in Control

ICML 2024oral

Temporal action abstractions, along with belief state representations, are a powerful knowledge sharing mechanism for sequential decision making. In this work, we propose a novel view that treats inducing temporal action abstractions as a sequence compression problem. To do so, we bring a subtle but…

2024

Premier-TACO is a Few-Shot Policy Learner: Pretraining Multitask Representation via Temporal Action-Driven Contrastive Loss

ICML 2024poster

We present Premier-TACO, a multitask feature representation learning approach designed to improve few-shot policy learning efficiency in sequential decision-making tasks. Premier-TACO leverages a subset of multitask offline datasets for pretraining a general feature representation, which captures cr…

2024

Successfully Guiding Humans with Imperfect Instructions by Highlighting Potential Errors and Suggesting Corrections

EMNLP 2024main

Language models will inevitably err in situations with which they are unfamiliar. However, by effectively communicating uncertainties, they can still guide humans toward making sound decisions in those contexts. We demonstrate this idea by developing HEAR, a system that can successfully guide humans…

2024

Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric Method

EMNLP 2024main

Extensive efforts in automated approaches for content moderation have been focused on developing models to identify toxic, offensive, and hateful content with the aim of lightening the load for moderators. Yet, it remains uncertain whether improvements on those tasks have truly addressed moderators’…

2024

Understanding the Impacts of Language Technologies’ Performance Disparities on African American Language Speakers

ACL 2024findings

This paper examines the experiences of African American Language (AAL) speakers when using language technologies. Previous work has used quantitative methods to uncover performance disparities between AAL speakers and White Mainstream English speakers when using language technologies, but has not so…

Cited by 12SourcePDFScholar
2024

“You Gotta be a Doctor, Lin” : An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations

EMNLP 2024main

Social science research has shown that candidates with names indicative of certain races or genders often face discrimination in employment practices. Similarly, Large Language Models (LLMs) have demonstrated racial and gender biases in various applications. In this study, we utilize GPT-3.5-Turbo a…

Cited by 0SourcePDFScholar
2023

$\texttt{TACO}$: Temporal Latent Action-Driven Contrastive Loss for Visual Reinforcement Learning

NeurIPS 2023poster

Despite recent progress in reinforcement learning (RL) from raw pixel data, sample inefficiency continues to present a substantial obstacle. Prior works have attempted to address this challenge by creating self-supervised auxiliary tasks, aiming to enrich the agent's learned representations with co…

2023

A Rose by Any Other Name would not Smell as Sweet: Social Bias in Names Mistranslation

EMNLP 2023long main

We ask the question: Are there widespread disparities in machine translations of names across race/ethnicity, and gender? We hypothesize that the translation quality of names and surrounding context will be lower for names associated with US racial and ethnic minorities due to these systems’ tendenc…

Cited by 0SourceScholar
2023

ASL Citizen: A Community-Sourced Dataset for Advancing Isolated Sign Language Recognition

NeurIPS 2023poster

Sign languages are used as a primary language by approximately 70 million D/deaf people world-wide. However, most communication technologies operate in spoken and written languages, creating inequities in access. To help tackle this problem, we release ASL Citizen, the first crowdsourced Isolated Si…

Cited by 35SourcePDFScholar
2023

Define, Evaluate, and Improve Task-Oriented Cognitive Capabilities for Instruction Generation Models

ACL 2023findings

Recent work studies the cognitive capabilities of language models through psychological tests designed for humans. While these studies are helpful for understanding the general capabilities of these models, there is no guarantee that a model possessing sufficient capabilities to pass those tests wou…

2023

Factual or Contextual? Disentangling Error Types in Entity Description Generation

ACL 2023long

In the task of entity description generation, given a context and a specified entity, a model must describe that entity correctly and in a contextually-relevant way. In this task, as well as broader language generation tasks, the generation of a nonfactual description (factual error) versus an incon…

2023

FairPrism: Evaluating Fairness-Related Harms in Text Generation

ACL 2023long

It is critical to measure and mitigate fairness-related harms caused by AI text generation systems, including stereotyping and demeaning harms. To that end, we introduce FairPrism, a dataset of 5,000 examples of AI-generated English text with detailed human annotations covering a diverse set of harm…

2023

It Takes Two to Tango: Navigating Conceptualizations of NLP Tasks and Measurements of Performance

ACL 2023findings

Progress in NLP is increasingly measured through benchmarks; hence, contextualizing progress requires understanding when and why practitioners may disagree about the validity of benchmarks. We develop a taxonomy of disagreement, drawing on tools from measurement modeling, and distinguish between two…

Cited by 16SourcePDFScholar
2023

Towards Conceptualization of ``Fair Explanation'': Disparate Impacts of anti-Asian Hate Speech Explanations on Content Moderators

EMNLP 2023long main

Recent research at the intersection of AI explainability and fairness has focused on how explanations can improve human-plus-AI task performance as assessed by fairness measures. We propose to characterize what constitutes an explanation that is itself "fair" -- an explanation that does not adversel…

Cited by 0SourcecodeScholar
2023

What Else Do I Need to Know? The Effect of Background Information on Users’ Reliance on QA Systems

EMNLP 2023long main

NLP systems have shown impressive performance at answering questions by retrieving relevant context. However, with the increasingly large models, it is impossible and often undesirable to constrain models' knowledge or reasoning to only the retrieved context. This leads to a mismatch between the inf…

Cited by 0SourceScholar
2023

Which Examples Should be Multiply Annotated? Active Learning When Annotators May Disagree

ACL 2023findings

Linguistic annotations, especially for controversial topics like hate speech detection, are frequently contested due to annotator backgrounds and positionalities. In such situations, preserving this disagreement through the machine learning pipeline can be important for downstream use cases. However…

2022

A Framework for Learning to Request Rich and Contextually Useful Information from Humans

ICML 2022spotlight

When deployed, AI agents will encounter problems that are beyond their autonomous problem-solving capabilities. Leveraging human assistance can help agents overcome their inherent limitations and robustly cope with unfamiliar situations. We present a general interactive framework that enables an age…

2022

Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications

NAACL 2022long

There are many ways to express similar things in text, which makes evaluating natural language generation (NLG) systems difficult. Compounding this difficulty is the need to assess varying quality criteria depending on the deployment setting. While the landscape of NLG evaluation has been well-mappe…

Cited by 36SourcePDFScholar
2022

Theory-Grounded Measurement of U.S. Social Stereotypes in English Language Models

NAACL 2022long

NLP models trained on text have been shown to reproduce human stereotypes, which can magnify harms to marginalized groups when systems are deployed at scale. We adapt the Agency-Belief-Communion (ABC) stereotype model of Koch et al. (2016) from social psychology as a framework for the systematic stu…

2021

A Novice-Reviewer Experiment to Address Scarcity of Qualified Reviewers in Large Conferences

AAAI 2021technical

Conference peer review constitutes a human-computation process whose importance cannot be overstated: not only it identifies the best submissions for acceptance, but, ultimately, it impacts the future of the whole research area by promoting some ideas and restraining others. A surge in the number of…

Cited by 41SourcePDFScholar
2021

Distantly-Supervised Dense Retrieval Enables Open-Domain Question Answering without Evidence Annotation

EMNLP 2021main

Open-domain question answering answers a question based on evidence retrieved from a large corpus. State-of-the-art neural approaches require intermediate evidence annotations for training. However, such intermediate annotations are expensive, and methods that rely on them cannot transfer to the mor…

2021

Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval

NAACL 2021long

Complex question answering often requires finding a reasoning chain that consists of multiple evidence pieces. Current approaches incorporate the strengths of structured knowledge and unstructured text, assuming text corpora is semi-structured. Building on dense retrieval methods, we propose a new m…

2019

Warm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback

ICML 2019oral

We investigate the feasibility of learning from both fully-labeled supervised data and contextual bandit data. We specifically consider settings in which the underlying learning signal may be different between these two data sources. Theoretically, we state and prove no-regret algorithms for learnin…

2018

Hierarchical Imitation and Reinforcement Learning

ICML 2018oral

We study how to effectively leverage expert feedback to learn sequential decision-making policies. We focus on problems with sparse rewards and long time horizons, which typically pose significant challenges in reinforcement learning. We propose an algorithmic framework, called hierarchical guidance…

Cited by 251SourcePDFScholar
2018

Residual Loss Prediction: Reinforcement Learning With No Incremental Feedback

ICLR 2018poster

We consider reinforcement learning and bandit structured prediction problems with very sparse loss feedback: only at the end of an episode. We introduce a novel algorithm, RESIDUAL LOSS PREDICTION (RESLOPE), that solves such problems by automatically learning an internal representation of a denser r…

Cited by 6SourcecodeScholar
2017

Active Learning for Cost-Sensitive Classification

ICML 2017poster

We design an active learning algorithm for cost-sensitive multiclass classification: problems where different errors have different costs. Our algorithm, COAL, makes predictions by regressing to each label’s cost and predicting the smallest. On a new example, it uses a set of regressors that perform…

Cited by 114SourcePDFScholar
2016

Opponent Modeling in Deep Reinforcement Learning

ICML 2016poster

Opponent modeling is necessary in multi-agent settings where secondary agents with competing goals also adapt their strategies, yet it remains challenging because of strategies’ complex interaction and the non-stationary nature. Most previous work focuses on developing probabilistic models or parame…

2015

Learning to Search Better than Your Teacher

ICML 2015poster

Methods for learning to search for structured prediction typically imitate a reference policy, with existing theoretical guarantees demonstrating low regret compared to that reference. This is unsatisfactory in many applications where the reference policy is suboptimal and the goal of learning is to…

Cited by 237SourcePDFScholar