Accelerating Linear Algebra Kernels on a Massively Parallel Reconfigurable Architecture
A. Soorishetty, Jian Zhou, Subhankar Pal, David T. Blaauw, H. Kim, Trevor N. Mudge, Ronald G. Dreslinski, Chaitali Chakrabarti
Abstract
Much of the recent work on domain-specific architectures has focused on bridging the gap between performance/efficiency and programmability. We consider one such example architecture, Transformer, consisting of light-weight cores interconnected by caches and crossbars that supports run-time reconfiguration between shared and private cache mode operations. We present customized implementation of a select set of linear algebra kernels, namely, triangular matrix solver, LU decomposition, QR decomposition and matrix in-version, on Transformer. The performance of the kernel algorithms is evaluated with respect to execution time and energy efficiency. Our study shows that each kernel achieves high performance for a certain cache mode and that this cache mode can change when the matrix size changes, making a case for run-time reconfiguration.
BibTeX
@inproceedings{icassp2020_acceleratingline,
title = {Accelerating Linear Algebra Kernels on a Massively Parallel Reconfigurable Architecture},
author = {A. Soorishetty and Jian Zhou and Subhankar Pal and David T. Blaauw and H. Kim and Trevor N. Mudge and Ronald G. Dreslinski and Chaitali Chakrabarti},
booktitle = {ICASSP 2020},
year = {2020}
}