← Search

Renhao Li

5 accepted papers

2025

CPsyExam: A Chinese Benchmark for Evaluating Psychology using Examinations

COLING 2025main

In this paper, we introduce a novel psychological benchmark, CPsyExam, constructed from questions sourced from Chinese examination systems. CPsyExam is designed to prioritize psychological knowledge and case analysis separately, recognizing the significance of applying psychological knowledge to rea…

2025

Exploring the Impact of Personality Traits on LLM Bias and Toxicity

EMNLP 2025

With the different roles that AI is expected to play in human life, imbuing large language models (LLMs) with different personalities has attracted increasing research interest. While the “personification” enhances human experiences of interactivity and adaptability of LLMs, it gives rise to critica

Cited by 0SourcePDFScholar
2025

HiMATE: A Hierarchical Multi-Agent Framework for Machine Translation Evaluation

EMNLP 2025

The advancement of Large Language Models (LLMs) enables flexible and interpretable automatic evaluations. In the field of machine translation evaluation, utilizing LLMs with translation error annotations based on Multidimensional Quality Metrics (MQM) yields more human-aligned judgments. However, cu

2024

CPsyCoun: A Report-based Multi-turn Dialogue Reconstruction and Evaluation Framework for Chinese Psychological Counseling

ACL 2024findings

Using large language models (LLMs) to assist psychological counseling is a significant but challenging task at present. Attempts have been made on improving empathetic conversations or acting as effective assistants in the treatment with LLMs. However, the existing datasets lack consulting knowledge…

2024

CoEvol: Constructing Better Responses for Instruction Finetuning through Multi-Agent Cooperation

EMNLP 2024main

In recent years, instruction fine-tuning (IFT) on large language models (LLMs) has garnered considerable attention to enhance model performance on unseen tasks. Attempts have been made on automatic construction and effective selection for IFT data. However, we posit that previous methods have not fu…