FlowNIB: An Information Bottleneck Analysis of Bidirectional vs. Unidirectional Language Models
Bidirectional language models (LMs) consistently show stronger context understanding than unidirectional models, yet the theoretical reason remains unclear. We present a simple information bottleneck (IB) perspective: bidirectional representations preserve more mutual information (MI) about both the…