← Search

Qiulin Zhang

4 accepted papers

2025

From Decoupling to Adaptive Transformation: a Wider Optimization Space for PTQ

ICLR 2025poster

Post-Training low-bit Quantization (PTQ) is useful to accelerate DNNs due to its high efficiency, the current SOTAs of which mostly adopt feature reconstruction with self-distillation finetuning. However, when bitwidth goes to be extremely low, we find the current reconstruction optimization space i…

Cited by 0SourcePDFScholar
2020

Model Rubik’s Cube: Twisting Resolution, Depth and Width for TinyNets

NeurIPS 2020poster

To obtain excellent deep neural architectures, a series of techniques are carefully designed in EfficientNets. The giant formula for simultaneously enlarging the resolution, depth and width provides us a Rubik’s cube for neural networks. So that we can find networks with high efficiency and excellen…

2020

Split to Be Slim: An Overlooked Redundancy in Vanilla Convolution

IJCAI 2020poster

Many effective solutions have been proposed to reduce the redundancy of models for inference acceleration. Nevertheless, common approaches mostly focus on eliminating less important filters or constructing efficient operations, while ignoring the pattern redundancy in feature maps. We reveal that ma…