← Search

Rohan Juneja

1 accepted papers

2026

HALO: Hardware-Aware Quantization with Low Critical-Path-Delay Weights for LLM Acceleration

AAAI 2026technical

Quantization is critical for efficiently deploying large language models (LLMs). Yet conventional methods remain hardware-agnostic, limited to bit-width constraints, and do not account for intrinsic circuit characteristics such as the timing behaviors and energy profiles of Multiply-Accumulate (MAC)

Cited by 0SourcePDFScholar