← Search

Tom Drummond

40 accepted papers

2026

Pixel-Level Residual Diffusion Transformer: Scalable 3D CT Volume Generation

ICLR 2026poster

Generating high-resolution 3D CT volumes with fine details remains challenging due to substantial computational demands and optimization difficulties inherent to existing generative models. In this paper, we propose the Pixel-Level Residual Diffusion Transformer (PRDiT), a scalable generative framew…

Cited by 0SourceScholar
2025

Robots Have Been Seen and Not Heard: Effects of Consequential Sounds on Human-Perception of Robots

RA-L 2025

Robots make compulsory machine sounds, known as “consequential sounds,” as they move and operate. As robots become more prevalent in workplaces, homes and public spaces, understanding how sounds produced by robots affect human-perceptions of these robots is becoming increasingly important to creatin

Cited by 6SourceScholar
2025

TCAM-Diff: Triplane-Aware Cross-Attention Medical Diffusion Model

AAAI 2025technical

We introduce TCAM-Diff, a novel 3D medical image generation model that reduces the memory requirements to encode and generate high-resolution 3D data. This model utilizes a decoder-only autoencoder method to learn triplane representation from dense volume and leverages generalization operations to p…

Cited by 0SourcePDFScholar
2024

Perceiving Longer Sequences With Bi-Directional Cross-Attention Transformers

NeurIPS 2024poster

We present a novel bi-directional Transformer architecture (BiXT) which scales linearly with input size in terms of computational cost and memory consumption, but does not suffer the drop in performance or limitation to only one input modality seen with other efficient Transformer-based approaches.…

2023

Boot and Switch: Alternating Distillation for Zero-Shot Dense Retrieval

EMNLP 2023long findings

Neural 'dense' retrieval models are state of the art for many datasets, however these models often exhibit limited domain transfer ability. Existing approaches to adaptation are unwieldy, such as requiring explicit supervision, complex model architectures, or massive external models. We present $\te…

Cited by 0SourcecodeScholar
2023

Knowledge Combination To Learn Rotated Detection Without Rotated Annotation

CVPR 2023poster

Rotated bounding boxes drastically reduce output ambiguity of elongated objects, making it superior to axis-aligned bounding boxes. Despite the effectiveness, rotated detectors are not widely employed. Annotating rotated bounding boxes is such a laborious process that they are not provided in many d…

2022

Adaptive Poincaré Point to Set Distance for Few-Shot Classification

AAAI 2022technical

Learning and generalizing from limited examples, i.e., few-shot learning, is of core importance to many real-world vision applications. A principal way of achieving few-shot learning is to realize an embedding where samples from different classes are distinctive. Recent studies suggest that embeddin…

Cited by 55SourcePDFScholar
2022

Implicit Motion Handling for Video Camouflaged Object Detection

CVPR 2022poster

We propose a new video camouflaged object detection (VCOD) framework that can exploit both short-term dynamics and long-term temporal consistency to detect camouflaged objects from video frames. An essential property of camouflaged objects is that they usually exhibit patterns similar to the backgro…

Cited by 104PDFcodeScholar
2022

Learning Instance and Task-Aware Dynamic Kernels for Few-Shot Learning

ECCV 2022poster

"Learning and generalizing to novel concepts with few samples (Few-Shot Learning) is still an essential challenge to real-world applications. A principle way of achieving few-shot learning is to realize a model that can rapidly adapt to the context of a given task. Dynamic networks have been shown c…

2022

On Enforcing Better Conditioned Meta-Learning for Rapid Few-Shot Adaptation

NeurIPS 2022accept

Inspired by the concept of preconditioning, we propose a novel method to increase adaptation speed for gradient-based meta-learning methods without incurring extra parameters. We demonstrate that recasting the optimisation problem to a non-linear least-squares formulation provides a principled way t…

Cited by 11SourcePDFScholar
2022

Rethinking Generalization in Few-Shot Classification

NeurIPS 2022accept

Single image-level annotations only correctly describe an often small subset of an image’s content, particularly when complex real-world scenes are depicted. While this might be acceptable in many classification scenarios, it poses a significant challenge for applications where the set of classes di…

2020

Hierarchical Neural Architecture Search for Deep Stereo Matching

NeurIPS 2020poster

To reduce the human efforts in neural network design, Neural Architecture Search (NAS) has been applied with remarkable success to various high-level vision tasks such as classification and semantic segmentation. The underlying idea for the NAS algorithm is straightforward, namely, to allow the netw…

2020

Learning to Take Good Pictures of People with a Robot Photographer

IROS 2020poster

We present a robotic system capable of navigating autonomously by following a line and taking good quality pictures of people. When a group of people is detected, the robot rotates towards them and then back to line while continuously taking pictures from different angles. Each picture is processed…

Cited by 11SourceScholar
2020

Reducing the Sim-to-Real Gap for Event Cameras

ECCV 2020poster

Event cameras are paradigm-shifting novel sensors that report asynchronous, per-pixel brightness changes called `events' with unparalleled low latency. This makes them ideal for high speed, high dynamic range scenes where conventional cameras would fail. Recent work has demonstrated impressive resul…

2020

Supportive Actions for Manipulation in Human-Robot Coworker Teams

IROS 2020poster

The increasing presence of robots alongside humans, such as in human-robot teams in manufacturing, gives rise to research questions about the kind of behaviors people prefer in their robot counterparts. We term actions that support interaction by reducing future interference with others as supportiv…

Cited by 9SourceScholar
2019

EMPNet: Neural Localisation and Mapping Using Embedded Memory Points

ICCV 2019poster

Continuously estimating an agent's state space and a representation of its surroundings has proven vital towards full autonomy. A shared common ground among systems which successfully achieve this feat is the integration of previously encountered observations into the current state being estimated.…

Cited by 24PDFScholar
2019

Event-Based Motion Segmentation by Motion Compensation

ICCV 2019poster

In contrast to traditional cameras, whose pixels have a common exposure time, event-based cameras are novel bio-inspired sensors whose pixels work independently and asynchronously output intensity changes (called "events"), with microsecond resolution. Since events are caused by the apparent motion…

Cited by 188PDFScholar
2019

LEARNING FACTORIZED REPRESENTATIONS FOR OPEN-SET DOMAIN ADAPTATION

ICLR 2019poster

Domain adaptation for visual recognition has undergone great progress in the past few years. Nevertheless, most existing methods work in the so-called closed-set scenario, assuming that the classes depicted by the target images are exactly the same as those of the source domain. In this paper, we ta…

Cited by 74SourcePDFScholar
2019

Look No Deeper: Recognizing Places from Opposing Viewpoints under Varying Scene Appearance using Single-View Depth Estimation

ICRA 2019poster

Visual place recognition (VPR) - the act of recognizing a familiar visual place - becomes difficult when there is extreme environmental appearance change or viewpoint change. Particularly challenging is the scenario where both phenomena occur simultaneously, such as when returning for the first time…

Cited by 29SourcecodeScholar
2019

Real-Time Joint Semantic Segmentation and Depth Estimation Using Asymmetric Annotations

ICRA 2019poster

Deployment of deep learning models in robotics as sensory information extractors can be a daunting task to handle, even using generic GPU cards. Here, we address three of its most prominent hurdles, namely, i) the adaptation of a single model to perform multiple tasks at once (in this work, we consi…

Cited by 169SourcecodeScholar
2019

The Importance of Metric Learning for Robotic Vision: Open Set Recognition and Active Learning

ICRA 2019poster

State-of-the-art deep neural network recognition systems are designed for a static and closed world. It is usually assumed that the distribution at test time will be the same as the distribution during training. As a result, classifiers are forced to categorise observations into one out of a set of…

Cited by 46SourceScholar
2018

CReaM: Condensed Real-time Models for Depth Prediction using Convolutional Neural Networks

IROS 2018poster

Since the resurgence of CNNs the robotic vision community has developed a range of algorithms that perform classification, semantic segmentation and structure prediction (depths, normals, surface curvature) using neural networks. While some of these models achieve state-of-the art results and super…

Cited by 23SourceScholar
2018

Efficient Subpixel Refinement With Symbolic Linear Predictors

CVPR 2018poster

We present an efficient subpixel refinement method using a learning-based approach called Linear Predictors. Firstly, we present a novel technique, called Symbolic Linear Predictors, which makes the learning step efficient for subpixel refinement. This makes our approach feasible for online applicat…

Cited by 1SourcePDFScholar
2018

Just-in-Time Reconstruction: Inpainting Sparse Maps Using Single View Depth Predictors as Priors

ICRA 2018poster

We present “just-in-time reconstruction” as realtime image-guided inpainting of a map with arbitrary scale and sparsity to generate a fully dense depth map for the image. In particular, our goal is to inpaint a sparse map - obtained from either a monocular visual SLAM system or a sparse sensor - usi…

Cited by 37SourceScholar
2017

Improved semantic segmentation for robotic applications with hierarchical conditional random fields

ICRA 2017poster

Conventional approaches to semantic segmentation are inappropriate for robotic applications, as they focus on pixel-level performance and give little significance to spurious object detections. This paper presents a region-based conditional random field model for semantic segmentation that focuses o…

Cited by 11SourceScholar
2017

Joint prediction of depths, normals and surface curvature from RGB images using CNNs

IROS 2017poster

Understanding the 3D structure of a scene is of vital importance, when it comes to developing fully autonomous robots. To this end, we present a novel deep learning based framework that estimates depth, surface normals and surface curvature by only using a single RGB image. To the best of our knowle…

Cited by 36SourceScholar
2016

MO-SLAM: Multi object SLAM with run-time object discovery through duplicates

IROS 2016poster

In this paper, we present MO-SLAM, a novel visual SLAM system that is capable of detecting duplicate objects in the scene during run-time without requiring an offline training stage to pre-populate a database of objects. Instead, we propose a novel method to detect landmarks that belong to duplicate…

Cited by 22SourceScholar
2015

Monocular image space tracking on a computationally limited MAV

ICRA 2015poster

We propose a method of monocular camera-inertial based navigation for computationally limited micro air vehicles (MAVs). Our approach is derived from the recent development of parallel tracking and mapping algorithms, but unlike previous results, we show how the tracking and mapping processes operat…

Cited by 8SourceScholar
2015

Self-calibration in visual sensor networks equipped with RGB-D cameras

ICASSP 2015accepted

We consider the self-calibration problem (estimation of location and orientation of multiple camera sensors), in visual sensor networks equipped with RGB-D cameras. We propose two algorithms based on feature matching and relative pose estimation. First one uses Floyd-Warshall algorithm, and can accu…

Cited by 0SourceScholar