← Search

Tonglan Xie

2 accepted papers

2026

Instruction-Guided Cross-Modal Clustering for Training-Free Visual Token Pruning in Vision-Language Models

AAAI 2026technical

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in understanding multimodal data such as images and text. However, the number of visual tokens in these models often far exceeds that of textual tokens, resulting in substantial redundancy and high inference costs. Existi

Cited by 0SourcePDFScholar
2025

Synergy Between the Strong and the Weak: Spiking Neural Networks are Inherently Self-Distillers

NeurIPS 2025poster

Brain-inspired spiking neural networks (SNNs) promise to be a low-power alternative to computationally intensive artificial neural networks (ANNs), although performance gaps persist. Recent studies have improved the performance of SNNs through knowledge distillation, but rely on large teacher models…

Cited by 0SourceScholar