2025
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
ICLR 2025poster
Large language models (LLMs) have showcased remarkable reasoning capabilities, yet they remain susceptible to errors, particularly in temporal reasoning tasks involving complex temporal logic. Existing research has explored LLM performance on temporal reasoning using diverse datasets and benchmarks.…