← Search

Bryan A. Plummer

40 accepted papers

2026

BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models

CVPR 2026

Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded framework for infant-inspired vision-language modeling that extensively improves upon BabyVLM-V1 through a longitudinal,

Cited by 0SourcecodeScholar
2026

CHAMMI-75: pre-training multi-channel models with heterogeneous microscopy images

ICLR 2026poster

Quantifying cell morphology using images and machine learning has proven to be a powerful tool to study the response of cells to treatments. However, the models used to quantify cellular morphology are typically trained with a single microscopy imaging type and under controlled experimental conditio…

Cited by 0SourcecodeScholar
2026

Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression

CVPR 2026

Parameter Recombination (PR) methods aim to efficiently compose the weights of a neural network, and encompasses tasks like Parameter-Efficient FineTuning (PEFT) and Model Compression (MC), among others. Most methods typically focus on one application of PR, which can make composing them challenging

Cited by 0SourceScholar
2026

Noise-Aware Generalization: Robustness to In-Domain Noise and Out-of-Domain Generalization

ICLR 2026poster

Methods addressing Learning with Noisy Labels (LNL) and multi-source Domain Generalization (DG) use training techniques to improve downstream task performance in the presence of label noise or domain shifts, respectively. Prior work often explores these tasks in isolation, with only limited work t…

Cited by 0SourcecodeScholar
2025

ChA-MAEViT: Unifying Channel-Aware Masked Autoencoders and Multi-Channel Vision Transformers for Improved Cross-Channel Learning

NeurIPS 2025poster

Prior work using Masked Autoencoders (MAEs) typically relies on random patch masking based on the assumption that images have significant redundancies across different channels, allowing for the reconstruction of masked content using cross-channel correlations. However, this assumption does not hold…

Cited by 0SourcecodeScholar
2025

Enhancing Virtual Try-On with Synthetic Pairs and Error-Aware Noise Scheduling

CVPR 2025poster

Given an isolated garment image in a canonical product view and a separate image of a person, the virtual try-on task aims to generate a new image of the person wearing the target garment.Prior virtual try-on works face two major challenges in achieving this goal: a) the paired (human, garment) trai…

Cited by 0SourcePDFScholar
2025

Is Large-scale Pretraining the Secret to Good Domain Generalization?

ICLR 2025poster

Multi-Source Domain Generalization (DG) is the task of training on multiple source domains and achieving high classification performance on unseen target domains. Recent methods combine robust features from web-scale pretrained backbones with new features learned from source data, and this has drama…

Cited by 1SourcePDFScholar
2025

Real, Fake, or Manipulated? Detecting Machine-Influenced Text

EMNLP 2025

Large Language Model (LLMs) can be used to write or modify documents, presenting a challenge for understanding the intent behind their use. For example, benign uses may involve using LLM on a human-written document to improve its grammar or to translate it into another language. However, a document

2025

Scaling Up Temporal Domain Generalization via Temporal Experts Averaging

EMNLP 2025

Temporal Domain Generalization (TDG) aims to generalize across temporal distribution shifts, e.g., lexical change over time. Prior work often addresses this by predicting future model weights. However, full model prediction is prohibitively expensive for even reasonably sized models. Thus, recent me

2025

Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning

EMNLP 2025

Large models achieve strong performance on Vision-and-Language Navigation (VLN) tasks, but are costly to run in resource-limited environments. Token pruning offers appealing tradeoffs for efficiency with minimal performance loss by reducing model input size, but prior work overlooks VLN-specific cha

2025

Web Artifact Attacks Disrupt Vision Language Models

ICCV 2025poster

Vision-language models (VLMs) (e.g., CLIP, LLaVA) are trained on large-scale, lightly curated web datasets, leading them to learn unintended correlations between semantic concepts and unrelated visual signals. These associations degrade model accuracy by causing predictions to rely on incidental pat…

2024

From Fake to Real: Pretraining on Balanced Synthetic Images to Prevent Spurious Correlations in Image Recognition

ECCV 2024oral

"Visual recognition models are prone to learning spurious correlations induced by a biased training set where certain conditions B (, Indoors) are over-represented in certain classes Y (, Big Dogs). Synthetic data from off-the-shelf large-scale generative models offers a promising direction to mitig…

2024

Koala: Key Frame-Conditioned Long Video-LLM

CVPR 2024highlight

Long video question answering is a challenging task that involves recognizing short-term activities and reasoning about their fine-grained relationships. State-of-the-art video Large Language Models (vLLMs) hold promise as a viable solution due to their demonstrated emergent capabilities on new task…

Cited by 33SourcePDFScholar
2024

Let Models Speak Ciphers: Multiagent Debate through Embeddings

ICLR 2024poster

Discussion and debate among Large Language Models (LLMs) have gained considerable attention due to their potential to enhance the reasoning ability of LLMs. Although natural language is an obvious choice for communication due to LLM's language understanding capability, the token sampling step needed…

Cited by 22SourcePDFScholar
2024

UniHuman: A Unified Model For Editing Human Images in the Wild

CVPR 2024poster

Human image editing includes tasks like changing a person's pose their clothing or editing the image according to a text prompt. However prior work often tackles these tasks separately overlooking the benefit of mutual reinforcement from learning them jointly. In this paper we propose UniHuman a uni…

2023

A Suite of Generative Tasks for Multi-Level Multimodal Webpage Understanding

EMNLP 2023long main

Webpages have been a rich, scalable resource for vision-language and language only tasks. Yet only pieces of webpages are kept in existing datasets: image-caption pairs, long text articles, or raw HTML, never all in one place. Webpage tasks have resultingly received little attention and structured i…

Cited by 0SourcecodeScholar
2023

Bias Mimicking: A Simple Sampling Approach for Bias Mitigation

CVPR 2023poster

Prior work has shown that Visual Recognition datasets frequently underrepresent bias groups B (e.g. Female) within class labels Y (e.g. Programmers). This dataset bias can lead to models that learn spurious correlations between class labels and bias groups such as age, gender, or race. Most recent m…

2023

CHAMMI: A benchmark for channel-adaptive models in microscopy imaging

NeurIPS 2023poster

Most neural networks assume that input images have a fixed number of channels (three for RGB images). However, there are many settings where the number of channels may vary, such as microscopy images where the number of channels changes depending on instruments and experimental goals. Yet, there has…

Cited by 11SourcePDFScholar
2023

Cola: A Benchmark for Compositional Text-to-image Retrieval

NeurIPS 2023poster

Compositional reasoning is a hallmark of human visual intelligence. Yet, despite the size of large vision-language models, they struggle to represent simple compositions by combining objects with their attributes. To measure this lack of compositional capability, we design Cola, a text-to-image retr…

Cited by 36SourcePDFScholar
2023

Collecting The Puzzle Pieces: Disentangled Self-Driven Human Pose Transfer by Permuting Textures

ICCV 2023poster

Human pose transfer synthesizes new view(s) of a person for a given pose. Recent work achieves this via self-reconstruction, which disentangles a person's pose and texture information by breaking down the person into several parts, then recombines them to reconstruct the person. However, this part-l…

Cited by 11PDFcodeScholar
2023

Language-Guided Audio-Visual Source Separation via Trimodal Consistency

CVPR 2023poster

We propose a self-supervised approach for learning to perform audio source separation in videos based on natural language queries, using only unlabeled video and audio pairs as training data. A key challenge in this task is learning to associate the linguistic description of a sound-emitting object…

Cited by 19SourcePDFScholar
2023

Show, Write, and Retrieve: Entity-aware Article Generation and Retrieval

EMNLP 2023long findings

Article comprehension is an important challenge in natural language processing with many applications such as article generation or image-to-article retrieval. Prior work typically encodes all tokens in articles uniformly using pretrained language models. However, in many applications, such as under…

Cited by 0SourcecodeScholar
2022

A Dataset for Interactive Vision-Language Navigation with Unknown Command Feasibility

ECCV 2022poster

"Vision-language navigation (VLN), in which an agent follows language instruction in a visual environment, has been studied under the premise that the input command is fully feasible in the environment. Yet in practice, a request may not be possible due to language ambiguity or environment changes.…

2022

Neural Parameter Allocation Search

ICLR 2022poster

Training neural networks requires increasing amounts of memory. Parameter sharing can reduce memory and communication costs, but existing methods assume networks have many identical layers and utilize hand-crafted sharing strategies that fail to generalize. We introduce Neural Parameter Allocation S…

2022

NewsStories: Illustrating Articles with Visual Summaries

ECCV 2022poster

"Recent self-supervised approaches have used large-scale image-text datasets to learn powerful representations that transfer to many tasks without finetuning. These methods often assume that there is one-to-one correspondence between its images and their (short) captions. However, many tasks require…

2022

Supervised Attribute Information Removal and Reconstruction for Image Manipulation

ECCV 2022poster

"The goal of attribute manipulation is to control specified attribute(s) in given images. Prior work approaches this problem by learning disentangled representations for each attribute that enables it to manipulate the encoded source attributes to the target attributes. However, encoded attributes a…

2021

CDS: Cross-Domain Self-Supervised Pre-Training

ICCV 2021poster

We present a two-stage pre-training approach that improves the generalization ability of standard single-domain pre-training. While standard pre-training on a single large dataset (such as ImageNet) can provide a good initial representation for transfer learning tasks, this approach may result in bi…

Cited by 57PDFScholar
2021

Effectively Leveraging Attributes for Visual Similarity

ICCV 2021poster

Measuring similarity between two images often requires performing complex reasoning along different axes (e.g., color, texture, or shape). Insights into what might be important for measuring similarity can can be provided by annotated attributes, but prior work tends to view these annotations as com…

Cited by 13PDFcodeScholar
2021

Look at What I’m Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos

NeurIPS 2021spotlight

We introduce the task of spatially localizing narrated interactions in videos. Key to our approach is the ability to learn to spatially localize interactions with self-supervision on a large corpus of videos with accompanying transcribed narrations. To achieve this goal, we propose a multilayer cro…

Cited by 26SourcePDFScholar
2020

Learning to Scale Multilingual Representations for Vision-Language Tasks

ECCV 2020poster

Current multilingual vision-language models either require a large number of additional parameters for each supported language, or suffer performance degradation as languages are added. In this paper, we propose a Scalable Multilingual Aligned Language Representation (SMALR) that supports many langu…

Cited by 37SourcePDFScholar
2020

Why do These Match? Explaining the Behavior of Image Similarity Models

ECCV 2020poster

Explaining a deep learning model can help users understand its behavior and allow researchers to discern its shortcomings. Recent work has primarily focused on explaining models for tasks like image classification or visual question answering. In this paper, we introduce Salient Attributes for Netwo…

2019

Language Features Matter: Effective Language Representations for Vision-Language Tasks

ICCV 2019poster

Shouldn't language and vision features be treated equally in vision-language (VL) tasks? Many VL approaches treat the language component as an afterthought, using simple language models that are either built upon fixed word embeddings trained on text-only data or are learned from scratch. We conclud…

Cited by 39PDFScholar
2019

Learning Similarity Conditions Without Explicit Supervision

ICCV 2019poster

Many real-world tasks require models to compare images along multiple similarity conditions (e.g. similarity in color, category or shape). Existing methods often reason about these complex similarity relationships by learning condition-aware embeddings. While such embeddings aid models in learning d…

Cited by 113PDFcodeScholar
2018

Conditional Image-Text Embedding Networks

ECCV 2018poster

This paper presents an approach for grounding phrases in images which jointly learns multiple text-conditioned embeddings in a single end-to-end model. In order to differentiate text phrases into semantically distinct subspaces, we propose a concept weight branch that automatically assigns phrases t…

2018

Learning Type-Aware Embeddings for Fashion Compatibility

ECCV 2018poster

Outfits in online fashion data are composed of items of many different types (e.g. top, bottom, shoes) that share some stylistic relationship with one another. A representation for building outfits requires a method that can learn both notions of similarity (for example, when two tops are interchang…

2017

Phrase Localization and Visual Relationship Detection With Comprehensive Image-Language Cues

ICCV 2017poster

This paper presents a framework for localization or grounding of phrases in images using a large collection of linguistic and visual cues. We model the appearance, size, and position of entity bounding boxes, adjectives that contain attribute information, and spatial relationships between pairs of e…

Cited by 230PDFcodeScholar
2015

Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

ICCV 2015poster

The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains linking mentions of the same entities in images, as well as 276k manually annotated boundi…

Cited by 2475PDFScholar