← Search

Shiai Zhu

5 accepted papers

2026

Exposing and Evaluating Hallucinations for GUI Grounding

CVPR 2026

Existing GUI benchmarks primarily focus on evaluating models' comprehensive capabilities but largely overlook hallucination phenomena in grounding tasks, which are crucial to the reliability of GUI understanding. In this work, we expose two major types of hallucinations in GUI grounding: 1) Confusio

Cited by 0SourceScholar
2026

QuietPrune: Query-Guided Early Token Pruning for Vision-Language Models

CVPR 2026

Vision-language models (VLMs) demonstrate powerful capabilities in multimodal tasks. However, the large number of visual tokens imposes a significant computational cost. In this paper, we propose QuietPrune, a QUery-guIded Early Token Pruning method to remove redundant visual tokens within VLMs, the

Cited by 0SourcecodeScholar
2026

SPEED-Q: Staged Processing with Enhanced Distillation Towards Efficient Low-Bit On-Device VLM Quantization

AAAI 2026technical

Deploying Vision-Language Models (VLMs) on edge devices (e.g., smartphones and robots) is crucial for enabling low-latency and privacy-preserving intelligent applications. Given the resource constraints of these devices, quantization offers a promising solution by improving memory efficiency and red

Cited by 0SourcePDFScholar
2025

LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades

NeurIPS 2025poster

As Large Language Models (LLMs) are frequently updated, LoRA weights trained on earlier versions quickly become obsolete. The conventional practice of retraining LoRA weights from scratch on the latest model is costly, time-consuming, and environmentally detrimental, particularly as the diversity of…

Cited by 0SourceScholar
2025

PQNAS: Mixed-precision Quantization-aware Neural Architecture Search with Pseudo Quantizer

ICASSP 2025accepted

Quantization-aware neural architecture search is an efficient way to automatically search for the best quantized model that can meet the limited resource constraints on edge devices. Existing methods utilize the straight-through estimator for training the quantized supernet, but lead to oscillation…

Cited by 0SourceScholar