← Search

Dongkun Shin

2 accepted papers

2026

FlexHiNM-GP: Flexible Hierarchical Pruning via Region Allocation and Channel Permutation

ICLR 2026poster

N:M sparsity has emerged as a hardware-friendly pruning strategy, notably supported by NVIDIA’s Sparse Tensor Cores. While efficient, its fixed sparsity ratio restricts flexibility, making it difficult to adapt pruning granularity to varying weight importance across layers and architectures. To over…

Cited by 0SourceScholar
2024

Proxyformer: Nyström-Based Linear Transformer with Trainable Proxy Tokens

AAAI 2024technical

Transformer-based models have demonstrated remarkable performance in various domains, including natural language processing, image processing and generative modeling. The most significant contributor to the successful performance of Transformer models is the self-attention mechanism, which allows fo…

Cited by 3SourcePDFScholar