← Search

Brandon McKinzie

4 accepted papers

2024

"MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training"

ECCV 2024poster

"In this work, we discuss building performant Multimodal Large Language Models (MLLMs). In particular, we study the importance of various architecture components and data choices. Through careful and comprehensive ablations of the image encoder, the vision language connector, and various pre-trainin…

2024

Pre-trained Language Models Do Not Help Auto-regressive Text-to-Image Generation

EMNLP 2024main

Recent advances in image tokenizers, such as VQ-VAE, have enabled text-to-image generation using auto-regressive methods, similar to language modeling. However, these methods have yet to leverage pre-trained language models, despite their adaptability to various downstream tasks. In this work, we ex…

Cited by 5SourcePDFScholar
2023

Perceptual Grouping in Contrastive Vision-Language Models

ICCV 2023poster

Recent advances in zero-shot image recognition suggest that vision-language models learn generic visual representations with a high degree of semantic information that may be arbitrarily probed with natural language phrases. Understanding an image, however, is not just about understanding what conte…

Cited by 53PDFScholar
2023

Robustness in Multimodal Learning under Train-Test Modality Mismatch

ICML 2023poster

Multimodal learning is defined as learning over multiple heterogeneous input modalities such as video, audio, and text. In this work, we are concerned with understanding how models behave as the type of modalities differ between training and deployment, a situation that naturally arises in many appl…

Cited by 6SourcePDFScholar