2025
Large language models can learn and generalize steganographic chain-of-thought under process supervision
NeurIPS 2025poster
Chain-of-thought (CoT) reasoning not only enhances large language model performance but also provides critical insights into decision-making processes, marking it as a useful tool for monitoring model intent and planning. By proactively preventing models from acting on CoT indicating misaligned or h…