← Search

Anthony Sicilia

12 accepted papers

2026

Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity

ICML 2026poster

Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collapse. Unlike prior work that attributes this effect to algorithmic limitations, we identify a fundamental, pervasive data-level driver: typicality bias in preference data, whereby annotators systematically…

Cited by 0SourceScholar
2025

Accounting for Sycophancy in Language Model Uncertainty Estimation

NAACL 2025findings

Effective human-machine collaboration requires machine learning models to externalize uncertainty, so users can reflect and intervene when necessary. For language models, these representations of uncertainty may be impacted by sycophancy bias: proclivity to agree with users, even if they are wrong.…

2025

Adaptive Platt Scaling with Causal Interpretations for Self-Reflective Language Model Uncertainty Estimates

EMNLP 2025

As large language models (LLMs) are consumed by more users and deployed in increasingly autonomous capacities, their ability to self-monitor and ask for human intervention is of vital importance. Underlying this capability are fundamental skills like self-reflection and expression of uncertainty. In

2025

Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others from Conversational Cues

ACL 2025long

Typically, when evaluating Theory of Mind, we consider the beliefs of others to be binary: held or not held. But what if someone is unsure about their own beliefs? How can we quantify this uncertainty? We propose a new suite of tasks, challenging language models (LMs) to model the uncertainty of par…

2025

Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation

EMNLP 2025

Establishing shared goals is a fundamental step in human-AI communication. However, ambiguities can lead to outputs that seem correct but fail to reflect the speaker’s intent. In this paper, we explore this issue with a focus on the data visualization domain, where ambiguities in natural language im

Cited by 0SourcePDFScholar
2025

SignAlignLM: Integrating Multimodal Sign Language Processing into Large Language Models

ACL 2025finding

Deaf and Hard-of-Hearing (DHH) users increasingly utilize Large Language Models (LLMs), yet face significant challenges due to these models’ limited understanding of sign language grammar, multimodal sign inputs, and Deaf cultural contexts. Further, current approaches that try to address these limit…

2024

Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models

ACL 2024findings

Effective interlocutors account for the uncertain goals, beliefs, and emotions of others. But even the best human conversationalist cannot perfectly anticipate the trajectory of a dialogue. How well can language models represent inherent uncertainty in conversations? We propose FortUne Dial, an expa…

2024

Generating Signed Language Instructions in Large-Scale Dialogue Systems

NAACL 2024industry

We introduce a goal-oriented conversational AI system enhanced with American Sign Language (ASL) instructions, presenting the first implementation of such a system on a worldwide multimodal conversational AI platform. Accessible through a touch-based interface, our system receives input from users a…

2023

Learning to Generate Equitable Text in Dialogue from Biased Training Data

ACL 2023long

The ingrained principles of fairness in a dialogue system’s decision-making process and generated responses are crucial for user engagement, satisfaction, and task achievement. Absence of equitable and inclusive principles can hinder the formation of common ground, which in turn negatively impacts t…

2022

PAC-Bayesian domain adaptation bounds for multiclass learners

UAI 2022poster

Multiclass neural networks are a common tool in modern unsupervised domain adaptation, yet an appropriate theoretical description for their non-uniform sample complexity is lacking in the adaptation literature. To fill this gap, we propose the first PAC-Bayesian adaptation bounds for multiclass lear…

2022

Test-time Fourier Style Calibration for Domain Generalization

IJCAI 2022poster

The topic of generalizing machine learning models learned on a collection of source domains to unknown target domains is challenging. While many domain generalization (DG) methods have achieved promising results, they primarily rely on the source domains at train-time without manipulating the target…

2022

The Change that Matters in Discourse Parsing: Estimating the Impact of Domain Shift on Parser Error

ACL 2022findings

Discourse analysis allows us to attain inferences of a text document that extend beyond the sentence-level. The current performance of discourse models is very low on texts outside of the training distribution’s coverage, diminishing the practical utility of existing models. There is need for a meas…