2024
Octopus: A Multi-modal LLM with Parallel Recognition and Sequential Understanding
NeurIPS 2024poster
A mainstream of Multi-modal Large Language Models (MLLMs) have two essential functions, i.e., visual recognition (e.g., grounding) and understanding (e.g., visual question answering). Presently, all these MLLMs integrate visual recognition and understanding in a same sequential manner in the LLM hea…