2025
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
NAACL 2025long
Large language models have demonstrated outstanding performance in various natural language processing tasks, but their security capabilities in the financial domain have not been explored, and their performance on complex tasks like financial agent remains unknown. This paper presents FinEval, a be…