← Search

Chaodong Xiao

3 accepted papers

2025

Polyline Path Masked Attention for Vision Transformer

NeurIPS 2025spotlight

Global dependency modeling and spatial position modeling are two core issues of the foundational architecture design in current deep learning frameworks. Recently, Vision Transformers (ViTs) have achieved remarkable success in computer vision, leveraging the powerful global dependency modeling capab…

Cited by 0SourcecodeScholar
2025

Spatial-Mamba: Effective Visual State Space Models via Structure-Aware State Fusion

ICLR 2025poster

Selective state space models (SSMs), such as Mamba, highly excel at capturing long-range dependencies in 1D sequential data, while their applications to 2D vision tasks still face challenges. Current visual SSMs often convert images into 1D sequences and employ various scanning patterns to incorpora…