AAAI 2026technical0 citations
FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research
Gabriel Recchia, Chatrik Singh Mangat, Issac Li, Gayatri Krishnakumar
Abstract
As AI models tackle increasingly complex problems, ensuring reliable human oversight becomes more challenging due to the difficulty of verifying solutions. Approaches to scaling AI supervision include debate, in which two agents engage in structured dialogue to help a judge evaluate claims; critique, in which models identify potential flaws in proposed solutions; and prover-verifier games, in which a capable
BibTeX
@inproceedings{aaai2026_findtheflawsanno,
title = {FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research},
author = {Gabriel Recchia and Chatrik Singh Mangat and Issac Li and Gayatri Krishnakumar},
booktitle = {AAAI 2026},
year = {2026}
}