← Search

Lian Yan

4 accepted papers

2026

AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models

AAAI 2026technical

n the agricultural domain, the deployment of large language models (LLMs) is hindered by the lack of training data and evaluation benchmarks. To mitigate this issue, we propose AgriEval, the first comprehensive Chinese agricultural benchmark with three main characteristics: (1) Comprehensive Capabil

Cited by 0SourcePDFScholar
2025

Agri-CM3: A Chinese Massive Multi-modal, Multi-level Benchmark for Agricultural Understanding and Reasoning

ACL 2025long

Multi-modal Large Language Models (MLLMs) integrating images, text, and speech can provide farmers with accurate diagnoses and treatment of pests and diseases, enhancing agricultural efficiency and sustainability. However, existing benchmarks lack comprehensive evaluations, particularly in multi-lev…

2025

KCS: Diversify Multi-hop Question Generation with Knowledge Composition Sampling

EMNLP 2025

Multi-hop question answering faces substantial challenges due to data sparsity, which increases the likelihood of language models learning spurious patterns. To address this issue, prior research has focused on diversifying question generation through content planning and varied expression. However,

2025

RLKGF: Reinforcement Learning from Knowledge Graph Feedback Without Human Annotations

ACL 2025finding

Reinforcement Learning from Human Feedback (RLHF) has been shown to effectively align large language models (LLMs) with human knowledge. However, the lack of human preference labels remains a significant bottleneck when applying RLHF to a downstream domain. Humans in RLHF play a critical role in inj…