← Search

Sadegh Mahdavi

5 accepted papers

2026

Scaling Generative Verifiers For Natural Language Mathematical Proof Verification And Selection

ICML 2026poster

Large language models have achieved remarkable success on final-answer mathematical problems, largely due to the ease of applying reinforcement learning with verifiable rewards. However, the reasoning underlying these solutions is often flawed. Advancing to rigorous proof-based mathematics requires …

Cited by 0SourceScholar
2025

Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation

ICML 2025poster

Advances in Large Language Models (LLMs) have sparked interest in their ability to solve Olympiad-level math problems. However, the training and evaluation of these models are constrained by the limited size and quality of available datasets, as creating large-scale data for such advanced problems…

2024

Leveraging Environment Interaction for Automated PDDL Translation and Planning with Large Language Models

NeurIPS 2024poster

Large Language Models (LLMs) have shown remarkable performance in various natural language tasks, but they often struggle with planning problems that require structured reasoning. To address this limitation, the conversion of planning problems into the Planning Domain Definition Language (PDDL) has…

2024

Memorization Capacity of Multi-Head Attention in Transformers

ICLR 2024spotlight

Transformers have become the go-to architecture for language and vision tasks, yet their theoretical properties, especially memorization capacity, remain elusive. This paper investigates the memorization abilities of multi-head attention mechanisms, examining how many example sequences they can memo…

2024

Revisiting the Equivalence of In-Context Learning and Gradient Descent: The Impact of Data Distribution

ICASSP 2024accepted

Transformers exhibit in-context learning (ICL), enabling adaptation to various tasks via prompts without the need for computationally intensive fine-tuning. Recent research investigates ICL’s mechanisms under analytically tractable models, with some conjecturing that ICL with linear attention implem…

Cited by 0SourceScholar