← Search

Shaowen Wang

6 accepted papers

2026

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

ICML 2026poster

Rotary Position Embeddings (RoPE) are widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads. Using simplified retrieval tasks and length generalization scenarios, we show—both empirica…

Cited by 0SourceScholar
2026

Deep Learning and Foundation Models for Weather Prediction: A Survey

IJCAI 2026

Numerical weather prediction (NWP) models remain the cornerstone of atmospheric sciences. Yet, deep learning (DL) is challenging this paradigm by its ability to capture intricate spatio-temporal patterns and deliver ultra-fast predictions. Analogous to the foundation models (e.g., ChatGPT) in natura

Cited by 0Scholar
2025

Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws

NeurIPS 2025spotlight

Large Language Models (LLMs) have demonstrated remarkable capabilities across numerous tasks, yet principled explanations for their underlying mechanisms and several phenomena, such as scaling laws, hallucinations, and related behaviors, remain elusive. In this work, we revisit the classical relatio…

Cited by 0SourceScholar
2023

Generative Table Pre-training Empowers Models for Tabular Prediction

EMNLP 2023long main

Recently, the topic of table pre-training has attracted considerable research interest. However, how to employ table pre-training to boost the performance of tabular prediction remains an open challenge. In this paper, we propose TapTap, the first attempt that leverages table pre-training to empower…

Cited by 0SourcecodeScholar