← Search

Michael Cogswell

8 accepted papers

2025

Punching Bag vs. Punching Person: Motion Transferability in Videos

ICCV 2025poster

Action recognition models demonstrate strong generalization, but can they effectively transfer high-level motion concepts across diverse contexts, even within similar distributions? For example, can a model recognize the broad action "punching" when presented with an unseen variation such as "punchi…

2024

BloomVQA: Assessing Hierarchical Multi-modal Comprehension

ACL 2024findings

We propose a novel VQA dataset, BloomVQA, to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. Unlike current benchmarks that often focus on fact-based memorization and simple reasoning tasks without theoretical grounding, we collect multiple-choice samples…

Cited by 0SourcePDFScholar
2024

DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback

CVPR 2024poster

We present DRESS a large vision language model (LVLM) that innovatively exploits Natural Language feedback (NLF) from Large Language Models to enhance its alignment and interactions by addressing two key limitations in the state-of-the-art LVLMs. First prior LVLMs generally rely only on the instruct…

Cited by 68SourcePDFScholar
2024

Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models

NAACL 2024long

Vision-language models (VLMs) have recently demonstrated strong efficacy as visual assistants that can parse natural queries about the visual content and generate human-like outputs. In this work, we explore the ability of these models to demonstrate human-like reasoning based on the perceived infor…

2022

Trigger Hunting with a Topological Prior for Trojan Detection

ICLR 2022poster

Despite their success and popularity, deep neural networks (DNNs) are vulnerable when facing backdoor attacks. This impedes their wider adoption, especially in mission critical applications. This paper tackles the problem of Trojan detection, namely, identifying Trojaned models – models trained with…

2020

Dialog without Dialog Data: Learning Visual Dialog Agents from VQA Data

NeurIPS 2020poster

Can we develop visually grounded dialog agents that can efficiently adapt to new tasks without forgetting how to talk to people? Such agents could leverage a larger variety of existing data to generalize to a new task, minimizing expensive data collection and annotation. In this work, we study a set…

2017

Grad-CAM: Visual Explanations From Deep Networks via Gradient-Based Localization

ICCV 2017poster

We propose a technique for producing 'visual explanations' for decisions from a large class of Convolutional Neural Network (CNN)-based models, making them more transparent. Our approach - Gradient-weighted Class Activation Mapping (Grad-CAM), uses the gradients of any target concept (say logits for…

Cited by 24144PDFcodeScholar
2016

Stochastic Multiple Choice Learning for Training Diverse Deep Ensembles

NeurIPS 2016poster

Many practical perception systems exist within larger processes which often include interactions with users or additional components that are capable of evaluating the quality of predicted solutions. In these contexts, it is beneficial to provide these oracle mechanisms with multiple highly likely h…

Cited by 233SourcePDFScholar