← Search

Jeffrey Wang

3 accepted papers

2026

Activation-Free Backbones for Image Recognition: Polynomial Alternatives for Spatial and Channel Mixing

ICML 2026poster

Modern vision backbones treat pointwise activations (e.g., ReLU, GELU) and exponential softmax as essential sources of nonlinearity, but we demonstrate they are not required. We design activation-free polynomial alternatives for three core primitives (MLPs, convolutions, and attention), where Hadama…

Cited by 0SourceScholar
2025

Position: Challenges and Future Directions of Data-Centric AI Alignment

ICML 2025poster

As AI systems become increasingly capable and influential, ensuring their alignment with human values, preferences, and goals has become a critical research focus. Current alignment methods primarily focus on designing algorithms and loss functions but often underestimate the crucial role of data. T…

Cited by 0SourcePDFScholar
2022

Efficient Large Scale Language Modeling with Mixtures of Experts

EMNLP 2022main

Mixture of Experts layers (MoEs) enable efficient scaling of language models through conditional computation. This paper presents a detailed empirical study of how autoregressive MoE language models scale in comparison with dense models in a wide range of settings: in- and out-of-domain language mod…

Cited by 146SourcecodeScholar