← Search

Anthony Dick

11 accepted papers

2024

BLURD: Benchmarking and Learning using a Unified Rendering and Diffusion Model

NeurIPS 2024poster

Recent advancements in pre-trained vision models have made them pivotal in computer vision, emphasizing the need for their thorough evaluation and benchmarking. This evaluation needs to consider various factors of variation, their potential biases, shortcuts, and inaccuracies that ultimately lead to…

2018

Visual Question Answering With Memory-Augmented Networks

CVPR 2018poster

In this paper, we exploit memory-augmented neural networks to predict accurate answers to visual questions, even when those answers rarely occur in the training set. The memory network incorporates both internal and external memory blocks and selectively pays attention to each training exemplar. We…

Cited by 134SourcePDFScholar
2017

DeepSetNet: Predicting Sets With Deep Neural Networks

ICCV 2017spotlight

This paper addresses the task of set prediction using deep learning. This is important because the output of many computer vision tasks, including image tagging and object detection, are naturally expressed as sets of entities rather than vectors. As opposed to a vector, the size of a set is not fix…

Cited by 56PDFScholar
2016

Ask Me Anything: Free-Form Visual Question Answering Based on Knowledge From External Sources

CVPR 2016spotlight

We propose a method for visual question answering which combines an internal representation of the content of an image with information extracted from a general knowledge base to answer a broad range of image-based questions. This allows more complex questions to be answered using the predominant ne…

Cited by 475PDFScholar
2016

Joint Probabilistic Matching Using m-Best Solutions

CVPR 2016oral

Matching between two sets of objects is typically approached by finding the object pairs that collectively maximize the joint matching score. In this paper, we argue that this single solution does not necessarily lead to the optimal matching accuracy and that general one-to-one assignment problems c…

Cited by 41PDFScholar
2016

Multi-modal Auto-Encoders as Joint Estimators for Robotics Scene Understanding

RSS 2016poster

We explore the capabilities of Auto-Encoders to fuse the information available from cameras and depth sensors, and to reconstruct missing data, for scene understanding tasks. In particular we consider three input modalities: RGB images; depth images; and semantic label information. We seek to genera…

Cited by 122SourcePDFScholar
2016

What Value Do Explicit High Level Concepts Have in Vision to Language Problems?

CVPR 2016poster

Much recent progress in Vision-to-Language (V2L) problems has been achieved through a combination of Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs). This approach does not explicitly represent high-level semantic concepts, but rather seeks to progress directly from image f…

Cited by 561PDFScholar
2015

A fast, modular scene understanding system using context-aware object detection

ICRA 2015poster

We propose a semantic scene understanding system that is suitable for real robotic operations. The system solves different tasks (semantic segmentation and object detections) in an opportunistic and distributed fashion but still allows communication between modules to improve their respective perfor…

Cited by 37SourceScholar
2015

Joint Probabilistic Data Association Revisited

ICCV 2015poster

In this paper, we revisit the joint probabilistic data association (JPDA) technique and propose a novel solution based on recent developments in finding the m-best solutions to an integer linear program. The key advantage of this approach is that it makes JPDA computationally tractable in applicatio…

Cited by 454PDFcodeScholar
2015

Part-Based Modelling of Compound Scenes From Images

CVPR 2015poster

We propose a method to recover the structure of a compound scene from multiple silhouettes. Structure is expressed as a collection of 3D primitives chosen from a pre-defined library, each with an associated pose. This has several advantages over a volume or mesh representation both for estimation an…

Cited by 33SourcePDFScholar