2025
Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
ICML 2025poster
Many existing evaluation benchmarks for Large Language Models (LLMs) quickly become outdated due to the emergence of new models and training data. These benchmarks also fall short in assessing how LLM performance changes over time, as they consist of a static set of questions without a temporal dime…