← Search

Yakoub Salhi

7 accepted papers

2026

Evaluating Robustness of Reasoning Models on Parameterized Logical Problems

ICML 2026oral

Logic provides a controlled testbed for evaluating LLM-based reasoners, yet standard SAT-style benchmarks often conflate surface difficulty (length, wording, clause order) with the structural phenomena that actually determine satisfiability. We introduce a diagnostic benchmark for \emph{2-SAT} built…

Cited by 0SourceScholar