2025
Semantic-KG: Using Knowledge Graphs to Construct Benchmarks for Measuring Semantic Similarity
NeurIPS 2025poster
Evaluating the open-form textual responses generated by Large Language Models (LLMs) typically requires measuring the semantic similarity of the response to a (human generated) reference. However, there is evidence that current semantic similarity methods may capture syntactic or lexical forms over…