Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
While large language-image pre-trained models like CLIP offer powerful generic features for image clustering, existing methods typically freeze the encoder. This creates a fundamental mismatch between the model