← Search

Linhao Yu

7 accepted papers

2025

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search

ACL 2025long

Video captioning can be used to assess the video understanding capabilities of Multimodal Large Language Models (MLLMs).However, existing benchmarks and evaluation protocols suffer from crucial issues, such as inadequate or homogeneous creation of key points, exorbitant cost of data creation, and li…

2025

Self-Pluralising Culture Alignment for Large Language Models

NAACL 2025long

As large language models (LLMs) become increasingly accessible in many countries, it is essential to align them to serve pluralistic human values across cultures. However, pluralistic culture alignment in LLMs remain an open problem. In this paper, we propose CultureSPA, a Self-Pluralising Culture A…

2025

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

ACL 2025long

Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing benchmarks for video understanding often treat these properties separately or narrowly focus on specific aspects, overlooki…

2024

CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models

ACL 2024findings

What a large language model (LLM) would respond in ethically relevant context? In this paper, we curate a large benchmark CMoralEval for morality evaluation of Chinese LLMs. The data sources of CMoralEval are two-fold: 1) a Chinese TV program discussing Chinese moral norms with stories from the soci…

2024

OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety

ACL 2024system demonstrations

The rapid development of Chinese large language models (LLMs) poses big challenges for efficient LLM evaluation. While current initiatives have introduced new benchmarks or evaluation platforms for assessing Chinese LLMs, many of these focus primarily on capabilities, usually overlooking potential a…

2023

CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion Types

EMNLP 2023long main

Spoken texts (either manual or automatic transcriptions from automatic speech recognition (ASR)) often contain disfluencies and grammatical errors, which pose tremendous challenges to downstream tasks. Converting spoken into written language is hence desirable. Unfortunately, the availability of da…

Cited by 0SourcecodeScholar