← Search

Hanqi Chen

3 accepted papers

2026

Enhancing Visual Token Representations for Video Large Language Models via Training-free Spatial-Temporal Pooling and Gridding

ICLR 2026poster

Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video understanding tasks, yet challenges remain in efficiently compressing visual tokens while preserving spatiotemporal interactions. Existing methods, such as LLaVA family, utilize simplistic pooling or interp…

Cited by 0SourcecodeScholar
2026

Sortblock: Similarity-Aware Feature Reuse for Diffusion Model

AAAI 2026technical

Diffusion Transformers (DiTs) have demonstrated remarkable generative capabilities, particularly benefiting from Transformer architectures that enhance visual and artistic fidelity. However, their inherently sequential denoising process results in high inference latency, limiting their deployment in

Cited by 0SourcePDFScholar
2025

Time-independent Spiking Neuron via Membrane Potential Estimation for Efficient Spiking Neural Networks

ICASSP 2025accepted

The computational inefficiency of spiking neural networks (SNNs) is primarily due to the sequential updates of membrane potential, which becomes more pronounced during extended encoding periods compared to artificial neural networks (ANNs). This highlights the need to parallelize SNN computations ef…

Cited by 0SourceScholar