← Search

Boao Kong

4 accepted papers

2026

CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure

ICLR 2026poster

Low-rank architectures have become increasingly important for efficient large language model (LLM) pre-training, providing substantial reductions in both parameter complexity and memory/computational demands. Despite these advantages, current low-rank methods face three critical shortcomings: (1) co…

Cited by 0SourceScholar
2026

Row-stochastic matrices can provably outperform doubly stochastic matrices in decentralized learning

ICML 2026poster

Decentralized learning often involves a weighted global loss with heterogeneous node weights $\lambda$. We revisit two natural strategies for incorporating these weights: (i) embedding them into the local losses to retain a uniform weight (and thus a doubly stochastic matrix), and (ii) keeping the o…

Cited by 0SourceScholar
2026

Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert Specialization

ICML 2026poster

Sparse Mixture-of-Experts (MoE) models scale Transformers efficiently but suffer from expert overlap, where different experts process similar tokens and learn redundant functions, resulting in ambiguous routing and underutilized capacity. While architectural solutions like DeepSeek-style shared expe…

Cited by 0SourceScholar
2024

SPARKLE: A Unified Single-Loop Primal-Dual Framework for Decentralized Bilevel Optimization

NeurIPS 2024poster

This paper studies decentralized bilevel optimization, in which multiple agents collaborate to solve problems involving nested optimization structures with neighborhood communications. Most existing literature primarily utilizes gradient tracking to mitigate the influence of data heterogeneity, with…

Cited by 2SourcePDFScholar