ICASSP 2019accepted0 citations

Cutensor-tubal: Optimized GPU Library for Low-tubal-rank Tensors

Tao Zhang, Xiao-Yang Liu

Abstract

In this paper, we optimize the computations of third-order low-tubal-rank tensor operations on many-core GPUs. Tensor operations are compute-intensive and existing studies optimize such operations in a case-by-case manner, which can be inefficient and error-prone. We develop and optimize a BLAS-like library for the low-tubal-rank tensor model called cuTensor-tubal, which includes efficient GPU primitives for tensor operations and key processes. We compute tensor operations in the frequency domain and fully exploit tube-wise and slice-wise parallelisms. We design, implement, and optimize four key tensor operations namely t-FFT, inverse t-FFT, t-product, and t-SVD. For t-product and t-SVD, cuTensor-tubal demonstrates significant speedups: maximum 29.16 ×, 6.72× speedups over the non-optimized GPU counterparts, and maximum 16.91× and 27.03× speedups over the CPU implementations running on dual 10-core Xeon CPUs.

BibTeX
@inproceedings{icassp2019_cutensortubalopt,
  title = {Cutensor-tubal: Optimized GPU Library for Low-tubal-rank Tensors},
  author = {Tao Zhang and Xiao-Yang Liu},
  booktitle = {ICASSP 2019},
  year = {2019}
}
Cutensor-tubal: Optimized GPU Library for Low-tubal-rank Tensors · ICASSP 2019