Truthfulness Does Not Scale Like Reasoning: Why Polling Fails as a Proxy Verifier
Pass@$k$ and other methods of scaling inference compute can improve language model performance in domains with external verifiers, including mathematics and code, where incorrect candidates can be filtered reliably. This raises a natural question: can we similarly scale compute to elicit gains in tr…