2024
SS1: Accelerating Inference with Fast and Expressive Sketch Structured Transform
NeurIPS 2024poster
Tensor multiplication with learned weight matrices is the fundamental building block in deep learning models. These matrices can often be sparsified, decomposed, quantized, or subjected to random parameter sharing without losing accuracy, suggesting the possibility of more efficient transforms. Alth…