← Search

Zhixuan Pan

2 accepted papers

2025

Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks

ICLR 2025poster

In this work, we investigate a particular implicit bias in gradient descent training, which we term “Feature Averaging,” and argue that it is one of the principal factors contributing to the non-robustness of deep neural networks. We show that, even when multiple discriminative features are present…

Cited by 1SourcePDFScholar
2025

Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws

NeurIPS 2025spotlight

Large Language Models (LLMs) have demonstrated remarkable capabilities across numerous tasks, yet principled explanations for their underlying mechanisms and several phenomena, such as scaling laws, hallucinations, and related behaviors, remain elusive. In this work, we revisit the classical relatio…

Cited by 0SourceScholar