← Search

Angela Zhou

15 accepted papers

2025

Fostering the Ecosystem of AI for Social Impact Requires Expanding and Strengthening Evaluation Standards

NeurIPS 2025poster

There has been increasing research interest in AI/ML for social impact, and correspondingly more publication venues refining review criteria for practice-driven AI/ML research. However, these review guidelines tend to most concretely recognize projects that simultaneously achieve deployment and nove…

Cited by 0SourceScholar
2021

It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks

NeurIPS 2021poster

Risk assessment instrument (RAI) datasets, particularly ProPublica’s COMPAS dataset, are commonly used in algorithmic fairness papers due to benchmarking practices of comparing algorithms on datasets used in prior work. In many cases, this data is used as a benchmark to demonstrate good performance…

Cited by 125SourceScholar
2019

Assessing Disparate Impact of Personalized Interventions: Identifiability and Bounds

NeurIPS 2019poster

Personalized interventions in social services, education, and healthcare leverage individual-level causal effect predictions in order to give the best treatment to each individual or to prioritize program interventions for the individuals most likely to benefit. While the sensitivity of these domain…

2019

Interval Estimation of Individual-Level Causal Effects Under Unobserved Confounding

AISTATS 2019poster

We study the problem of learning conditional average treatment effects (CATE) from observational data with unobserved confounders. The CATE function maps baseline covariates to individual causal effect predictions and is key for personalized assessments. Recent work has focused on how to learn CATE…

Cited by 122SourcePDFScholar
2019

The Fairness of Risk Scores Beyond Classification: Bipartite Ranking and the XAUC Metric

NeurIPS 2019poster

Where machine-learned predictive risk scores inform high-stakes decisions, such as bail and sentencing in criminal justice, fairness has been a serious concern. Recent work has characterized the disparate impact that such risk scores can have when used for a binary classification task. This may not…