← Search

Sejoon Oh

1 accepted papers

2024

Cross-Modal Projection in Multimodal LLMs Doesn’t Really Project Visual Attributes to Textual Space

ACL 2024short

Multimodal large language models (MLLMs) like LLaVA and GPT-4(V) enable general-purpose conversations about images with the language modality. As off-the-shelf MLLMs may have limited capabilities on images from domains like dermatology and agriculture, they must be fine-tuned to unlock domain-specif…