← Search

Tess Wood

3 accepted papers

2025

Incorporating Diverse Perspectives in Cultural Alignment: Survey of Evaluation Benchmarks Through A Three-Dimensional Framework

EMNLP 2025

Large Language Models (LLMs) increasingly serve diverse global audiences, making it critical for responsible AI deployment across cultures. While recent works have proposed various approaches to enhance cultural alignment in LLMs, a systematic analysis of their evaluation benchmarks remains needed.

Cited by 0SourcePDFScholar
2025

SEEval: Advancing LLM Text Evaluation Efficiency and Accuracy through Self-Explanation Prompting

NAACL 2025findings

Large language models (LLMs) have achieved remarkable success in various natural language generation (NLG) tasks, but their performance in automatic text evaluation is not yet ready as human replacements. In this paper, we propose SEEval (Self-Explanation in Evaluation), a novel prompt-based text ev…

Cited by 0SourcePDFScholar
2024

HalluMeasure: Fine-grained Hallucination Measurement Using Chain-of-Thought Reasoning

EMNLP 2024main

Automating the measurement of hallucinations in LLM generated responses is a challenging task as it requires careful investigation of each factual claim in a response. In this paper, we introduce HalluMeasure, a new LLM-based hallucination detection mechanism that decomposes an LLM response into ato…

Cited by 3SourcePDFScholar