One Patch Doesn’t Fit All: Adaptive Patching for Native-Resolution Multimodal Large Language Models
Real-world visual signals are inherently variable in resolution, and it is natural to endow multimodal large language models (MLLMs) with such native-resolution perception capabilities. In principle, for general and straightforward multimodal understanding, low-resolution images are sufficient. Whil…