← Search

Shuzi Niu

5 accepted papers

2026

FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix Sketching

AAAI 2026technical

Traditional post-training quantization (PTQ) is considered an effective approach to reduce model size and accelerate inference of large-scale language models (LLMs). However, existing low-rank PTQ methods require costly fine-tuning to determine a compromise rank for diverse data and layers in large

Cited by 0SourcePDFScholar
2026

Factorization-in-Loop:Proximal Fill-in Minimization for Sparse Matrix Reordering

AAAI 2026technical

Fill-ins are new nonzero elements in the summation of the upper and lower triangular factors generated during LU factorization. For large sparse matrices, they will increase the memory usage and computational time, and be reduced through proper row or column arrangement, namely matrix reordering. Fi

Cited by 0SourcePDFScholar
2026

Learning Fill-in Reduction Ordering via Graph Policy Optimization for Sparse Matrices

ICASSP 2026poster

Matrix reordering in large sparse solvers seeks a permutation that minimizes factorization fill-in to reduce memory and computation. Because the minimum fill-in ordering problem is NP-complete and fill-in is implicit in the sparsity pattern, graph-theoretic heuristics are used. Existing reinforcemen…

Cited by 0SourcePDFScholar