← Search

Daniel Lo

2 accepted papers

2023

Dynamic Stashing Quantization for Efficient Transformer Training

EMNLP 2023short findings

Large Language Models (LLMs) have demonstrated impressive performance on a range of Natural Language Processing (NLP) tasks. Unfortunately, the immense amount of computations and memory accesses required for LLM training makes them prohibitively expensive in terms of hardware cost, and thus challeng…

Cited by 0SourceScholar
2020

Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating Point

NeurIPS 2020poster

In this paper, we explore the limits of Microsoft Floating Point (MSFP), a new class of datatypes developed for production cloud-scale inferencing on custom hardware. Through the co-evolution of hardware design and algorithms, MSFP achieves accuracy comparable to or better than industry standards Bf…