2026
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
ICML 2026poster
Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain unclear. Chain-of-thoughts (CoTs) have been commonly utilized to externalize LLMs' thinking, but this strategy fails to accurately reflect LLMs' thinking process.…