2025
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
NeurIPS 2025poster
Recent advances in reasoning techniques have substantially improved the performance of large language models (LLMs), raising expectations for their ability to provide accurate, truthful, and reliable information. However, emerging evidence suggests that iterative reasoning may foster belief entrench…