← Search

Young Geun Kim

2 accepted papers

2025

First Attentions Last: Better Exploiting First Attentions for Efficient Parallel Training

NeurIPS 2025poster

As training billion-scale transformers becomes increasingly common, employing multiple distributed GPUs along with parallel training methods has become a standard practice. However, existing transformer designs suffer from significant communication overhead, especially in Tensor Parallelism (TP), wh…

Cited by 0SourceScholar
2025

Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models

NAACL 2025long

Transformer-based large-scale pre-trained models achieve great success. Fine-tuning is the standard practice for leveraging these models in downstream tasks. Among the fine-tuning methods, adapter-tuning provides a parameter-efficient fine-tuning by introducing lightweight trainable modules while ke…

Cited by 0SourcePDFScholar