ICASSP 2025accepted0 citations

FG3DFormer: Fine-Grained 3D Shape Classification Based on Vision Transformer

Xiangyu Ma, Jing Bai, Jinzhe Jiang, Bin Peng

Abstract

Fine-grained 3D shape classification (FGSC) remains challenging due to the difficulty of adaptively capturing global structure differences and subtle inter-class distinctions. This paper directly extends Vision Transformer (ViT) to FGSC, proposing a pure Transformer network FG3DFormer that fully leverages ViT’s global correlation and local attention abilities. FG3Dformer comprises the Hierarchical Feature Extraction (HFE) and the Hierarchical Feature Refinement (HFR), interconnected through the Adaptive View Region Selection (AVRS). Firstly, the HFE comprehensively evaluates the significance of intra-view patches and views driven by inter-view and intraview attention. Then, the AVRS adaptively selects crucial patch Tokens from different views to serve as sources of subtle local features. Finally, the HFR refines the 3D shape descriptor, capturing more discriminative global and subtle local features by leveraging both the view and selected crucial patch Tokens. Extensive experiments on FG3D and ModelNet40 demonstrate the superiority of FG3Dformer in FGSC and meta-category 3D shape classification tasks.

BibTeX
@inproceedings{icassp2025_fg3dformerfinegr,
  title = {FG3DFormer: Fine-Grained 3D Shape Classification Based on Vision Transformer},
  author = {Xiangyu Ma and Jing Bai and Jinzhe Jiang and Bin Peng},
  booktitle = {ICASSP 2025},
  year = {2025}
}
FG3DFormer: Fine-Grained 3D Shape Classification Based on Vision Transformer · ICASSP 2025