2025
Charm: The Missing Piece in ViT Fine-Tuning for Image Aesthetic Assessment
CVPR 2025poster
The capacity of Vision transformers (ViTs) to handle variable-sized inputs is often constrained by computational complexity and batch processing limitations. Consequently, ViTs are typically trained on small, fixed-size images obtained through downscaling or cropping. While reducing computational bu…