← Search

Christopher M. Homan

9 accepted papers

2026

Forest vs Tree: The (N, K) Trade-off in Reproducible ML Evaluation

AAAI 2026technical

Reproducibility is a cornerstone of scientific validation and of the authority it confers on its results. Reproducibility in machine learning evaluations leads to greater trust, confidence, and value. However, the ground truth responses used in machine learning often necessarily come from humans, am

Cited by 0SourcePDFScholar
2026

ProRefine: Inference-Time Prompt Refinement with Textual Feedback (Student Abstract)

AAAI 2026technical

Agentic workflows, where multiple AI agents collaborate to accomplish complex tasks like reasoning or planning, play a substantial role in many cutting-edge commercial applications. These workflows depend critically on the prompts used to provide the roles models play in such workflows. Poorly desig

Cited by 0SourcePDFScholar
2025

ARTICLE: Annotator Reliability Through In-Context Learning

AAAI 2025technical

Ensuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsically subjective, creating a challenging scenario for traditional quality assessment approaches because it is hard to dist…

2025

ARTICLE: Annotator Reliability Through In-Context Learning (Student Abstract)

AAAI 2025technical

Ensuring annotator quality in training and evaluation data is a key piece of machine learning in NLP. Tasks such as sentiment analysis and offensive speech detection are intrinsically subjective, creating a challenging scenario for traditional quality assessment approaches because it is hard to dist…

2025

Bayelemabaga: Creating Resources for Bambara NLP

NAACL 2025long

Data curation for under-resource languages enables the development of more accurate and culturally sensitive natural language processing models. However, the scarcity of well-structured multilingual datasets remains a challenge for advancing machine translation in these languages, especially for Afr…

Cited by 0SourcePDFScholar
2025

GAIfE: Using GenAI to Improve Literacy in Low-resourced Settings

NAACL 2025findings

Illiteracy is a predictor of many negative social and personal outcomes. Illiteracy rates are particularly high in countries with underresourced languages, where few books exist that are suitable for children to learn to read from. We present GAIfE (Generative AI for Education), a toolchain and work…

2025

Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope Speech

EMNLP 2025

This paper makes three contributions. First, via a substantial corpus of 1,419,047 comments posted on 3,161 YouTube news videos of major US cable news outlets, we analyze how users engage with LGBTQ+ news content. Our analyses focus both on positive and negative content. In particular, we construct

2024

Rater Cohesion and Quality from a Vicarious Perspective

EMNLP 2024finding

Human feedback is essential for building human-centered AI systems across domains where disagreement is prevalent, such as AI safety, content moderation, or sentiment analysis. Many disagreements, particularly in politically charged settings, arise because raters have opposing values or beliefs. Vic…

2023

Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive

EMNLP 2023long main

Offensive speech detection is a key component of content moderation. However, what is offensive can be highly subjective. This paper investigates how machine and human moderators disagree on what is offensive when it comes to real-world social web political discourse. We show that (1) there is exten…

Cited by 0SourcecodeScholar