2026
Statistically Undetectable Backdoors in Deep Neural Networks
ICML 2026poster
We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the f…