← Search

Erick Draayer

1 accepted papers

2025

Benchmarking Large Language Models with Integer Sequence Generation Tasks

NeurIPS 2025poster

We present a novel benchmark designed to rigorously evaluate the capabilities of large language models (LLMs) in mathematical reasoning and algorithmic code synthesis tasks. The benchmark comprises integer sequence generation tasks sourced from the Online Encyclopedia of Integer Sequences (OEIS), te…

Cited by 0SourceScholar