← Search

Jaemin Park

2 accepted papers

2025

First Attentions Last: Better Exploiting First Attentions for Efficient Parallel Training

NeurIPS 2025poster

As training billion-scale transformers becomes increasingly common, employing multiple distributed GPUs along with parallel training methods has become a standard practice. However, existing transformer designs suffer from significant communication overhead, especially in Tensor Parallelism (TP), wh…

Cited by 0SourceScholar
2025

From Theory to Practice: Rethinking Green and Martin Kernels for Unleashing Graph Transformers

ICML 2025poster

Graph Transformers (GTs) have emerged as a powerful alternative to message-passing neural networks, yet their performance heavily depends on effectively embedding structural inductive biases. In this work, we introduce novel structural encodings (SEs) grounded in a rigorous analysis of random walks…

Cited by 0SourcePDFScholar