← Search

Haoqi Zhu

4 accepted papers

2026

SMM Transformer: Leveraging Spiking Neural Networks for Multimodal Tasks

ICML 2026poster

Spiking Neural Networks (SNNs) enable event-driven computation with sparse activations, but building multimodal Transformers on SNNs is hindered by unstable training in deep spiking stacks and a mismatch between dense softmax attention and spike-based communication. We propose SMM Transformer, an SN…

Cited by 0SourceScholar
2026

Spike-HTR: Spiking Neural Transformer for Handwritten Text Recognition

ICML 2026poster

Offline handwritten text recognition (HTR) is blank-dominated: task-relevant evidence lies in sparse ink strokes, yet mainstream recognizers still expend dense spatial compute and full-length width-axis token mixing across the canvas. Spiking neural networks (SNNs) promise activity-proportional comp…

Cited by 0SourceScholar
2024

HaltingVT: Adaptive Token Halting Transformer for Efficient Video Recognition

ICASSP 2024accepted

Action recognition in videos poses a challenge due to its high computational cost, especially for Joint Space-Time video transformers (Joint VT). Despite their effectiveness, the excessive number of tokens in such architectures significantly limits their efficiency. In this paper, we propose Halting…

Cited by 0SourceScholar
2020

Multi-Way Multi-View Deep Autoencoder for Image Feature Learning with Multi-Level Graph Regularization

ICASSP 2020accepted

Multi-view feature learning has garnered much attention recently since many real world data are comprised of different representations or views. How to explore the consensus structure and eliminate the inconsistency noise in different views remains a challenging problem in multi-view feature learnin…

Cited by 0SourceScholar