← Search

Malgorzata Gwiazda

1 accepted papers

2026

TimeSeriesExamAgent: Creating TimeSeries Reasoning Benchmarks at Scale

ICLR 2026poster

Large Language Models (LLMs) have shown promising performance in time series modeling tasks, but do they truly understand time series data? While multiple benchmarks have been proposed to answer this fundamental question, most are manually curated and focus on narrow domains or specific skill sets.…

Cited by 0SourcecodeScholar