2024
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
ECCV 2024poster
"We all know that large language models are built on top of a transformer-based architecture to process textual inputs. For example, the LLaMA family of models stands out among many open-source implementations. Can the same transformer be used to process 2D images? In this paper, we answer this ques…