← Search

Paul Newman

56 accepted papers

2026

Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference

ICLR 2026poster

Intuitive physics understanding in video diffusion models plays an essential role in building general-purpose physically plausible world simulators, yet accurately evaluating such capacity remains a challenging task due to the difficulty in disentangling physics correctness from visual appearance in…

Cited by 0SourcecodeScholar
2025

Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models

NeurIPS 2025poster

Foundation models (FMs) trained with different objectives and data learn diverse representations, making some more effective than others for specific downstream tasks. Existing adaptation strategies, such as parameter-efficient fine-tuning, focus on individual models and do not exploit the complemen…

Cited by 0SourceScholar
2024

Masked γ-SSL: Learning Uncertainty Estimation via Masked Image Modeling

ICRA 2024poster

This work proposes a semantic segmentation network that produces high-quality uncertainty estimates in a single forward pass. We exploit general representations from foundation models and unlabelled datasets through a Masked Image Modeling (MIM) approach, which is robust to augmentation hyper-parame…

Cited by 1SourceScholar
2024

NeuralFloors++: Consistent Street-Level Scene Generation From BEV Semantic Maps

IROS 2024poster

Learning autonomous driving capabilities requires diverse and realistic training data. This has led to exploring generative techniques as an alternative to real-world data collection. In this paper we propose a method for synthesising photo-realistic urban driving scenes, along with semantic, instan…

Cited by 0SourceScholar
2024

NeuralFloors: Conditional Street-Level Scene Generation From BEV Semantic Maps via Neural Fields

RA-L 2024

Semantic Bird's Eye View (BEV) representations are a popular format, being easily interpretable and editable. However, synthesising ground-view images from BEVs is a difficult task as the system would need to learn both the mapping from BEV to Front View (FV) structure as well as to synthesise highl

Cited by 2SourceScholar
2024

RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Multi-Modal Large Language Model Learning

RSS 2024poster

We need to trust robots that use often opaque AI methods. They need to explain themselves to us, and we need to trust their explanation. In this regard, explainability plays a critical role in trustworthy autonomous decision-making to foster transparency and acceptance among end users, especially in…

Cited by 83SourcePDFScholar
2024

That’s My Point: Compact Object-centric LiDAR Pose Estimation for Large-scale Outdoor Localisation

ICRA 2024poster

This paper is about 3D pose estimation on LiDAR scans with extremely minimal storage requirements to enable scalable mapping and localisation. We achieve this by clustering all points of segmented scans into semantic objects and representing them only with their respective centroid and semantic clas…

Cited by 3SourceScholar
2024

VDNA-PR: Using General Dataset Representations for Robust Sequential Visual Place Recognition

ICRA 2024poster

This paper adapts a general dataset representation technique to produce robust Visual Place Recognition (VPR) descriptors, crucial to enable real-world mobile robot localisation. Two parallel lines of work on VPR have shown, on one side, that general-purpose off-the-shelf feature representations can…

Cited by 1SourceScholar
2023

Off the Radar: Uncertainty-Aware Radar Place Recognition with Introspective Querying and Map Maintenance

IROS 2023poster

Localisation with Frequency-Modulated Continuous-Wave (FMCW) radar has gained increasing interest due to its inherent resistance to challenging environments. However, complex artefacts of the radar measurement process require appropriate uncertainty estimation - to ensure the safe and reliable appli…

Cited by 8SourceScholar
2023

Visual DNA: Representing and Comparing Images Using Distributions of Neuron Activations

CVPR 2023poster

Selecting appropriate datasets is critical in modern computer vision. However, no general-purpose tools exist to evaluate the extent to which two datasets differ. For this, we propose representing images -- and by extension datasets -- using Distributions of Neuron Activations (DNAs). DNAs fit distr…

Cited by 13SourcePDFScholar
2023

Visual Servoing on Wheels: Robust Robot Orientation Estimation in Remote Viewpoint Control

IROS 2023poster

This work proposes a fast deployment pipeline for visually-servoed robots which does not assume anything about either the robot - e.g. sizes, colour or the presence of markers - or the deployment environment. Specifically, we apply a learning based approach to reliably estimate the pose of a robot i…

Cited by 5SourceScholar
2022

BoxGraph: Semantic Place Recognition and Pose Estimation from 3D LiDAR

IROS 2022poster

This paper is about extremely robust and lightweight localisation using LiDAR point clouds based on instance segmentation and graph matching. We model 3D point clouds as fully-connected graphs of semantically identified components where each vertex corresponds to an object instance and encodes its s…

Cited by 27SourceScholar
2022

Fast-MbyM: Leveraging Translational Invariance of the Fourier Transform for Efficient and Accurate Radar Odometry

ICRA 2022poster

Masking by Moving (MByM), provides robust and accurate radar odometry measurements through an exhaustive correlative search across discretised pose candidates. However, this dense search creates a significant computational bottleneck which hinders real-time performance when high-end GPUs are not ava…

Cited by 31SourceScholar
2022

What Goes Around: Leveraging a Constant-Curvature Motion Constraint in Radar Odometry

RA-L 2022

This letter presents a method that leverages vehicle motion constraints to refine data associations in a point-based radar odometry system. By using the strong prior on how a non-holonomic robot is constrained to move smoothly through its environment, we develop the necessary framework to estimate e

Cited by 18SourceScholar
2021

Fool Me Once: Robust Selective Segmentation via Out-of-Distribution Detection with Contrastive Learning

ICRA 2021poster

In this work, a neural network is trained to simultaneously perform segmentation and pixel-wise Out-of-Distribution (OoD) detection, such that the segmentation of unknown regions of scenes can be rejected. This is made possible by leveraging an OoD dataset with a novel contrastive objective and data…

Cited by 16SourceScholar
2021

Get to the Point: Learning Lidar Place Recognition and Metric Localisation Using Overhead Imagery

RSS 2021poster

This paper is about localising a robot in overhead images using lidar. Specifically; we show how to solve both place recognition and metric localisation of a lidar using only publicly available overhead imagery as a map proxy. This is in contrast to current approaches that rely on prior sensor map…

Cited by 29SourcePDFScholar
2020

Kidnapped Radar: Topological Radar Localisation using Rotationally-Invariant Metric Learning

ICRA 2020poster

This paper presents a system for robust, large-scale topological localisation using Frequency-Modulated Continuous-Wave scanning radar which extends the state-of-the-art by an efficient, learning-based approach to handle radar data for localisation. We learn a metric space for embedding polar radar…

Cited by 85SourceScholar
2020

Self-Supervised Localisation between Range Sensors and Overhead Imagery

RSS 2020poster

Publicly available satellite imagery can be an ubiquitous, cheap, and powerful tool for vehicle localisation when a prior sensor map is unavailable. However, satellite images are not directly comparable to data from ground range sensors because of their starkly different modalities. We present a l…

Cited by 26SourcePDFScholar
2020

The Oxford Radar RobotCar Dataset: A Radar Extension to the Oxford RobotCar Dataset

ICRA 2020poster

In this paper we present The Oxford Radar RobotCar Dataset, a new dataset for researching scene understanding using Millimetre-Wave FMCW scanning radar data. The target application is autonomous vehicles where this modality is robust to environmental conditions such as fog, rain, snow, or lens flare…

Cited by 494SourceScholar
2019

Fast Radar Motion Estimation with a Learnt Focus of Attention using Weak Supervision

ICRA 2019poster

This paper is about fast motion estimation with scanning radar. We use weak supervision to train a focus of attention policy which actively down-samples the measurement stream before data association steps are undertaken. At training, we avoid laborious manual labelling by exploiting short-term sens…

Cited by 69SourceScholar
2018

Adversarial Training for Adverse Conditions: Robust Metric Localisation Using Appearance Transfer

ICRA 2018poster

We present a method of improving visual place recognition and metric localisation under very strong appearance change. We learn an invertable generator that can transform the conditions of images, e.g. from day to night, summer to winter etc. This image transforming filter is explicitly designed to…

Cited by 125SourceScholar
2018

Fast Global Labelling for Depth-Map Improvement Via Architectural Priors

ICRA 2018poster

Depth map estimation techniques from cameras often struggle to accurately estimate the depth of large textureless regions. In this work we present a vision-only method that accurately extracts planar priors from a viewed scene without making any assumptions of the underlying scene layout. Through a…

Cited by 1SourceScholar
2018

Geometric Multi-Model Fitting With a Convex Relaxation Algorithm

CVPR 2018poster

We propose a novel method for fitting multiple geometric models to multi-structural data via convex relaxation. Unlike greedy methods - which maximise the number of inliers - our approach efficiently searches for a soft assignment of points to geometric models by minimising the energy of the overall…

Cited by 39SourcePDFScholar
2018

Mark Yourself: Road Marking Segmentation via Weakly-Supervised Annotations from Multimodal Data

ICRA 2018poster

This paper presents a weakly-supervised learning system for real-time road marking detection using images of complex urban environments obtained from a monocular camera. We avoid expensive manual labelling by exploiting additional sensor modalities to generate large quantities of annotated images in…

Cited by 45SourceScholar
2018

Multimotion Visual Odometry (MVO): Simultaneous Estimation of Camera and Third-Party Motions

IROS 2018poster

Estimating motion from images is a well-studied problem in computer vision and robotics. Previous work has developed techniques to estimate the motion of a moving camera in a largely static environment (e.g., visual odometry) and to segment or track motions in a dynamic scene using known camera moti…

Cited by 60SourceScholar
2018

Precise Ego-Motion Estimation with Millimeter-Wave Radar Under Diverse and Challenging Conditions

ICRA 2018poster

In contrast to cameras, lidars, GPS, and proprioceptive sensors, radars are affordable and efficient systems that operate well under variable weather and lighting conditions, require no external infrastructure, and detect long-range objects. In this paper, we present a reliable and accurate radar-on…

Cited by 232SourceScholar
2018

Resource-Performance Tradeoff Analysis for Mobile Robots

RA-L 2018

The design of mobile autonomous robots is challenging due to the limited on-board resources such as processing power and energy. A promising approach is to generate intelligent schedules that reduce the resource consumption while maintaining best performance, or more interestingly, to tradeoff reduc

Cited by 33SourceScholar
2018

Surface Edge Explorer (see): Planning Next Best Views Directly from 3D Observations

ICRA 2018poster

Surveying 3D scenes is a common task in robotics. Systems can do so autonomously by iteratively obtaining measurements. This process of planning observations to improve the model of a scene is called Next Best View (NBV) planning. NBV planning approaches often use either volumetric (e.g., voxel grid…

Cited by 45SourceScholar
2017

NID-SLAM: Robust Monocular SLAM Using Normalised Information Distance

CVPR 2017poster

We propose a direct monocular SLAM algorithm based on the Normalised Information Distance (NID) metric. In contrast to current state-of-the-art direct methods based on photometric error minimisation, our information-theoretic NID metric provides robustness to appearance variation due to lighting, we…

Cited by 74PDFcodeScholar
2016

A unified representation for application of architectural constraints in large-scale mapping

ICRA 2016

This paper is about discovering and leveraging architectural constraints in large scale 3D reconstructions using laser. Our contribution is to offer a formulation of the problem which naturally and in a unified way, captures the variety of architectural constraints that can be discovered and applied

Cited by 1SourceScholar
2016

Choosing a time and place for calibration of lidar-camera systems

ICRA 2016

We propose a calibration method that automatically estimates the extrinsic calibration between a sensor pose-graph from natural scenes. The sensor pose-graph represents a system of sensors comprising of lidars and cameras, without sensor co-visibility constraints. The method addresses the fact that

Cited by 32SourceScholar
2016

Made to measure: Bespoke landmarks for 24-hour, all-weather localisation with a camera

ICRA 2016

This paper is about camera-only localisation in challenging outdoor environments, where changes in lighting, weather and season cause traditional localisation systems to fail. Conventional approaches to the localisation problem rely on point-features such as SIFT, SURF or BRIEF to associate landmark

Cited by 88SourceScholar
2016

The path less taken: A fast variational approach for scene segmentation used for closed loop control

IROS 2016poster

In this paper we propose an on-line system that discovers and drives collision-free traversable paths, using a variational approach to dense stereo vision. Our system is light weight, can be run on low cost hardware and is remarkably quick to predict the semantics. In addition to the scene's path af…

Cited by 9SourceScholar
2015

A variational approach to online road and path segmentation with monocular vision

ICRA 2015poster

In this paper we present an online approach to segmenting roads on large scale trajectories using only a monocular camera mounted on a car. We differ from popular 2D segmentation solutions which use single colour images and machine learning algorithms that require supervised training on huge image d…

Cited by 14SourceScholar
2015

Exploiting known unknowns: Scene induced cross-calibration of lidar-stereo systems

IROS 2015poster

We propose an automatic, targetless, data-driven, extrinsic calibration method to calibrate push-broom 2D lidars with a multi-camera system. The calibration problem is decoupled into alternating optimisers over two hierarchical levels, where both levels are linked with a penalty term. The lower-leve…

Cited by 25SourceScholar
2015

FARLAP: Fast robust localisation using appearance priors

ICRA 2015poster

This paper is concerned with large-scale localisation at city scales with monocular cameras. Our primary motivation lies with the development of autonomous road vehicles - an application domain in which low-cost sensing is particularly important. Here we present a method for localising against a tex…

Cited by 47SourceScholar
2015

From dusk till dawn: Localisation at night using artificial light sources

ICRA 2015poster

This paper is about localising at night in urban environments using vision. Despite it being dark exactly half of the time, surprisingly little attention has been given to this problem. A defining aspect of night-time urban scenes is the presence and effect of artificial lighting - be that in the fo…

Cited by 42SourceScholar
2015

Integrating metric and semantic maps for vision-only automated parking

ICRA 2015poster

We present a framework for integrating two layers of map which are often required for fully automated operation: metric and semantic. Metric maps are likely to improve with subsequent visitations to the same place, while semantic maps can comprise both permanent and fluctuating features of the envir…

Cited by 34SourceScholar
2015

Know your limits: Embedding localiser performance models in teach and repeat maps

ICRA 2015poster

This paper is about building maps which not only contain the traditional information useful for localising — such as point features — but also embeds a spatial model of expected localiser performance. This often overlooked second-order information provides vital context when it comes to map use and…

Cited by 24SourceScholar
2015

Too much TV is bad: Dense reconstruction from sparse laser with non-convex regularisation

ICRA 2015poster

In this paper we address the problem of dense depth map estimation from sparse noisy range data to reconstruct large heterogeneous outdoor scenes. We propose a surface inpainting solution through energy minimisation with an adaptive selection of surface regularisers among a set of well known convex…

Cited by 19SourceScholar
2015

Work smart, not hard: Recalling relevant experiences for vast-scale but time-constrained localisation

ICRA 2015poster

This paper is about life-long vast-scale localisation in spite of changes in weather, lighting and scene structure. Building upon our previous work in Experience-based Navigation [1], we continually grow and curate a visual map of the world that explicitly supports multiple representations of the sa…

Cited by 137SourceScholar