← Search

Sarah Tan

3 accepted papers

2025

Evaluating Cultural and Social Awareness of LLM Web Agents

NAACL 2025findings

As large language models (LLMs) expand into performing as agents for real-world applications beyond traditional NLP tasks, evaluating their robustness becomes increasingly important. However, existing benchmarks often overlook critical dimensions like cultural and social awareness. To address these,…

2023

Error Discovery By Clustering Influence Embeddings

NeurIPS 2023poster

We present a method for identifying groups of test examples---slices---on which a model under-performs, a task now known as slice discovery. We formalize coherence---a requirement that erroneous predictions, within a slice, should be wrong for the same reason---as a key property that any slice disco…

Cited by 4SourcePDFScholar
2020

Purifying Interaction Effects with the Functional ANOVA: An Efficient Algorithm for Recovering Identifiable Additive Models

AISTATS 2020poster

Models which estimate main effects of individual variables alongside interaction effects have an identifiability challenge: effects can be freely moved between main effects and interaction effects without changing the model prediction. This is a critical problem for interpretability because it permi…