← Search

Yuhang Cai

3 accepted papers

2026

LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws

ICML 2026poster

Existing scaling laws for Large Language Models (LLMs), predominantly monotonic power laws, have successfully guided model development but fail to explain emerging non-monotonic phenomena such as catastrophic overtraining and quantization-induced degradation, where performance deteriorates despite i…

Cited by 0SourceScholar
2025

Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks

ICML 2025poster

We establish the asymptotic implicit bias of gradient descent (GD) for generic non-homogeneous deep networks under exponential loss. Specifically, we characterize three key properties of GD iterates starting from a sufficiently small empirical risk, where the threshold is determined by a measure of…

Cited by 0SourcePDFScholar
2024

Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization

NeurIPS 2024poster

The typical training of neural networks using large stepsize gradient descent (GD) under the logistic loss often involves two distinct phases, where the empirical risk oscillates in the first phase but decreases monotonically in the second phase. We investigate this phenomenon in two-layer networks…

Cited by 6SourcePDFScholar