← Search

Tianruo Rose Xu

2 accepted papers

2026

Can You Trust What I Think? Analyzing and Improving Verbalized Uncertainty and Factuality in Reasoning-Based Large Language Models

AAAI 2026technical

Reasoning-based large language models now often produce natural-language thinking traces alongside their answers, but it remains unclear whether these verbalized uncertainties faithfully reflect their knowledge or can be used to improve factuality. We study this question for long-form, knowledge-int

Cited by 0SourcePDFScholar
2025

The Progress Illusion: Revisiting meta-evaluation standards of LLM evaluators

EMNLP 2025

LLM judges have gained popularity as an inexpensive and performant substitute for human evaluation. However, we observe that the meta-evaluation setting in which the reliability of these LLM evaluators is established is substantially different from their use in model development. To address this, we