← Search

Xinfeng Xie

2 accepted papers

2024

CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor Optimization

NeurIPS 2024poster

Training large AI models such as LLMs and DLRMs costs massive GPUs and computing time. The high training cost has become only affordable to big tech companies, meanwhile also causing increasing concerns about the environmental impact. This paper presents CoMERA, a **Co**mputing- and **M**emory-**E**…

2018

HitNet: Hybrid Ternary Recurrent Neural Network

NeurIPS 2018poster

Quantization is a promising technique to reduce the model size, memory footprint, and massive computation operations of recurrent neural networks (RNNs) for embedded devices with limited resources. Although extreme low-bit quantization has achieved impressive success on convolutional neural networks…

Cited by 76SourcePDFScholar