2026
Maximizing mutual information between prompt and response improves LLM performance with no additional data
ICML 2026poster
While post-training has successfully improved large language models across a variety of domains from open-ended text generation to mathematics, these gains heavily rely on human-labeled data or external verifiers. Existing data has already been exploited and new high-quality data is expensive to col…