2026
Causal Interpretation of Neural Network Computations with Contribution Decomposition (CODEC)
ICLR 2026poster
Understanding how neural networks transform inputs into outputs is crucial for interpreting and manipulating their behavior. Most existing approaches analyze internal representations by identifying hidden-layer activation patterns correlated with human-interpretable concepts. Here we take a direct a…