AAAI 2026technical0 citations

Points Meet Pixels: Bridging 2D Vision-Language Model and 3D Perception Gaps for Point Cloud Quality Assessment

Mingxuan Li, Zihao Huang, Xiaohui Chu, Fazhan Zhang, Bohan Fu, Runze Hu

Abstract

Vision-Language Models (VLMs) have demonstrated significant progress in quality assessment tasks. However, a fundamental paradox arises when their application to Point Cloud Quality Assessment (PCQA). Existing VLMs, designed for image-text pairs, are inherently incompatible with 3D point cloud data due to the modality gap. While some PCQA research attempts to adapt point clouds to VLMs by 2D projection, this approach inevitably sacrifices crucial spatial structure information essential for accurate quality assessment. Conversely, directly integrating a dedicated 3D branch into a VLM-based PCQA framework introduces feature space misalignment and an influx of quality-insensitive information. To bridge these fundamental conflicts hindering VLMs

BibTeX
@inproceedings{aaai2026_pointsmeetpixels,
  title = {Points Meet Pixels: Bridging 2D Vision-Language Model and 3D Perception Gaps for Point Cloud Quality Assessment},
  author = {Mingxuan Li and Zihao Huang and Xiaohui Chu and Fazhan Zhang and Bohan Fu and Runze Hu},
  booktitle = {AAAI 2026},
  year = {2026}
}
Points Meet Pixels: Bridging 2D Vision-Language Model and 3D Perception Gaps for Point Cloud Quality Assessment · AAAI 2026