← Search

Wenjing Ke

2 accepted papers

2024

DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models

EMNLP 2024industry

Improving the efficiency of inference in Large Language Models (LLMs) is a critical area of research. Post-training Quantization (PTQ) is a popular technique, but it often faces challenges at low-bit levels, particularly in downstream tasks. Quantization-aware Training (QAT) can alleviate this probl…

Cited by 2SourcePDFScholar
2022

Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-Identification

CVPR 2022poster

Recently, self-attention mechanisms have shown impressive performance in various NLP and CV tasks, which can help capture sequential characteristics and derive global information. In this work, we explore how to extend self-attention modules to better learn subtle feature embeddings for recognizing…

Cited by 230PDFScholar