← Search

Keren Zhou

2 accepted papers

2024

SS1: Accelerating Inference with Fast and Expressive Sketch Structured Transform

NeurIPS 2024poster

Tensor multiplication with learned weight matrices is the fundamental building block in deep learning models. These matrices can often be sparsified, decomposed, quantized, or subjected to random parameter sharing without losing accuracy, suggesting the possibility of more efficient transforms. Alth…

2023

Hardware-Aware Compression with Random Operation Access Specific Tile (ROAST) Hashing

ICML 2023poster

Advancements in deep learning are often associated with increasing model sizes. Training and deploying large models require sophisticated hardware and incur significantly higher costs. Thus, model compression is a widely explored approach to solving the problem. However, SOTA techniques fall short i…