← Search

Avishree Khare

1 accepted papers

2025

Logicbreaks: A Framework for Understanding Subversion of Rule-based Inference

ICLR 2025poster

We study how to subvert large language models (LLMs) from following prompt-specified rules. We first formalize rule-following as inference in propositional Horn logic, a mathematical system in which rules have the form "if $P$ and $Q$, then $R$" for some propositions $P$, $Q$, and $R$. Next, we prov…

Cited by 1SourcePDFScholar