← Search

Daniel O'Malley

2 accepted papers

2025

Benchmarking Large Language Models with Integer Sequence Generation Tasks

NeurIPS 2025poster

We present a novel benchmark designed to rigorously evaluate the capabilities of large language models (LLMs) in mathematical reasoning and algorithmic code synthesis tasks. The benchmark comprises integer sequence generation tasks sourced from the Online Encyclopedia of Integer Sequences (OEIS), te…

Cited by 0SourceScholar
2025

Model-Agnostic Knowledge Guided Correction for Improved Neural Surrogate Rollout

ICLR 2025poster

Modeling the evolution of physical systems is critical to many applications in science and engineering. As the evolution of these systems is governed by partial differential equations (PDEs), there are a number of computational simulations which resolve these systems with high accuracy. However, as…