← Search

Hongyaoxing Gu

3 accepted papers

2026

FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix Sketching

AAAI 2026technical

Traditional post-training quantization (PTQ) is considered an effective approach to reduce model size and accelerate inference of large-scale language models (LLMs). However, existing low-rank PTQ methods require costly fine-tuning to determine a compromise rank for diverse data and layers in large

Cited by 0SourcePDFScholar
2026

TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D Tiling

ICML 2026poster

Mixture-of-Experts (MoE) models achieve remarkable performance by sparsely activating specialized experts, yet their massive parameters in experts pose significant challenges for deployment. While low-rank quantization offers a promising route to compress MoE models, existing methods still incur non…

Cited by 0SourceScholar
2024

A differentiable brain simulator bridging brain simulation and brain-inspired computing

ICLR 2024poster

Brain simulation builds dynamical models to mimic the structure and functions of the brain, while brain-inspired computing (BIC) develops intelligent systems by learning from the structure and functions of the brain. The two fields are intertwined and should share a common programming framework to f…

Cited by 4SourcePDFScholar