2024
Large Language Models' Expert-level Global History Knowledge Benchmark (HiST-LLM)
NeurIPS 2024poster
Large Language Models (LLMs) have the potential to transform humanities and social science research, yet their history knowledge and comprehension at a graduate level remains untested. Benchmarking LLMs in history is particularly challenging, given that human knowledge of history is inherently unbal…