← Search

YI XIAODIE

1 accepted papers

2026

FlexHiNM-GP: Flexible Hierarchical Pruning via Region Allocation and Channel Permutation

ICLR 2026poster

N:M sparsity has emerged as a hardware-friendly pruning strategy, notably supported by NVIDIA’s Sparse Tensor Cores. While efficient, its fixed sparsity ratio restricts flexibility, making it difficult to adapt pruning granularity to varying weight importance across layers and architectures. To over…

Cited by 0SourceScholar