← Search

Richard Bowden

42 accepted papers

2025

Addressing Dimensional Scaling in Reinforcement Learning for Symbolic Locomotion Policies through Leveraging Inductive Priors

IROS 2025

We explore symbolic policy optimization for various legged locomotion challenges; specifically walker environments ranging from bipedal to highly redundant systems with 128 legs. These represent a broad range of action space dimensionalities. We find that state-of-the-art symbolic policy optimizatio

Cited by 0SourceScholar
2025

Geo-Sign: Hyperbolic Contrastive Regularisation for Geometrically Aware Sign Language Translation

NeurIPS 2025poster

Recent progress in Sign Language Translation has focussed primarily on improving the representational capacity of large language models to incorporate sign-language features. This work explores an alternative direction: enhancing the geometric properties of skeletal representations themselves. We pr…

Cited by 0SourcecodeScholar
2025

The Radiance of Neural Fields: Democratizing Photorealistic and Dynamic Robotic Simulation

ICRA 2025

As robots increasingly coexist with humans, they must navigate complex, dynamic environments rich in visual information and implicit social dynamics, like when to yield or move through crowds. Addressing these challenges requires significant advances in vision-based sensing and a deeper understandin

Cited by 0SourceScholar
2024

Campus Map: A Large-Scale Dataset to Support Multi-View VO, SLAM and BEV Estimation

ICRA 2024poster

Significant advances in robotics and machine learning have resulted in many datasets designed to support research into autonomous vehicle technology. However, these datasets are rarely suitable for a wide variety of navigation tasks. For example, datasets that include multiple cameras often have sho…

Cited by 1SourceScholar
2024

Select and Reorder: A Novel Approach for Neural Sign Language Production

COLING 2024main

Sign languages, often categorised as low-resource languages, face significant challenges in achieving accurate translation due to the scarcity of parallel annotated datasets. This paper introduces Select and Reorder (S&R), a novel approach that addresses data scarcity by breaking down the translatio…

Cited by 2SourcePDFScholar
2024

Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language Translation

ICLR 2024poster

Automatic Sign Language Translation requires the integration of both computer vision and natural language processing to effectively bridge the communication gap between sign and spoken languages. However, the deficiency in large-scale training data to support sign language translation means we need…

Cited by 36SourcePDFScholar
2023

Kick Back & Relax: Learning to Reconstruct the World by Watching SlowTV

ICCV 2023poster

Self-supervised monocular depth estimation (SS-MDE) has the potential to scale to vast quantities of data. Unfortunately, existing approaches limit themselves to the automotive domain, resulting in models incapable of generalizing to complex environments such as natural or indoor settings. To addres…

Cited by 21PDFcodeScholar
2022

"The Pedestrian Next to the Lamppost" Adaptive Object Graphs for Better Instantaneous Mapping

CVPR 2022poster

Estimating a semantically segmented bird's-eye-view (BEV) map from a single image has become a popular technique for autonomous control and navigation. However, they show an increase in localization error with distance from the camera. While such an increase in error is entirely expected - localizat…

Cited by 8PDFScholar
2022

AFT-VO: Asynchronous Fusion Transformers for Multi-View Visual Odometry Estimation

IROS 2022poster

Motion estimation approaches typically employ sensor fusion techniques, such as the Kalman Filter, to handle individual sensor failures. More recently, deep learning-based fusion approaches have been proposed, increasing the performance and requiring less model-specific implementations. However, cur…

Cited by 9SourceScholar
2022

BEV-SLAM: Building a Globally-Consistent World Map Using Monocular Vision

IROS 2022poster

The ability to produce large-scale maps for nav-igation, path planning and other tasks is a crucial step for autonomous agents, but has always been challenging. In this work, we introduce BEV-SLAM, a novel type of graph-based SLAM that aligns semantically-segmented Bird's Eye View (BEV) predictions…

Cited by 11SourceScholar
2022

Learning an Interpretable Model for Driver Behavior Prediction with Inductive Biases

IROS 2022poster

To plan safe maneuvers and act with foresight, autonomous vehicles must be capable of accurately predicting the uncertain future. In the context of autonomous driving, deep neural networks have been successfully applied to learning pre-dictive models of human driving behavior from data. However, the…

Cited by 8SourcecodeScholar
2022

Signing at Scale: Learning to Co-Articulate Signs for Large-Scale Photo-Realistic Sign Language Production

CVPR 2022poster

Sign languages are visual languages, with vocabularies as rich as their spoken language counterparts. However, current deep-learning based Sign Language Production (SLP) models produce under-articulated skeleton pose sequences from constrained vocabularies and this limits applicability. To be unders…

Cited by 76PDFcodeScholar
2021

Enabling spatio-temporal aggregation in Birds-Eye-View Vehicle Estimation

ICRA 2021poster

Constructing Birds-Eye-View (BEV) maps from monocular images is typically a complex multi-stage process involving the separate vision tasks of ground plane estimation, road segmentation and 3D object detection. However, recent approaches have adopted end-to-end solutions which warp image-based featu…

Cited by 67SourceScholar
2021

Markov Localisation using Heatmap Regression and Deep Convolutional Odometry

ICRA 2021poster

In the context of self-driving vehicles there is strong competition between approaches based on visual localisation and Light Detection And Ranging (LiDAR). While LiDAR provides important depth information, it is sparse in resolution and expensive. On the other hand, cameras are low-cost and recent…

Cited by 1SourceScholar
2021

Mixed SIGNals: Sign Language Production via a Mixture of Motion Primitives

ICCV 2021poster

It is common practice to represent spoken languages at their phonetic level. However, for sign languages, this implies breaking motion into its constituent motion primitives. Avatar based Sign Language Production (SLP) has traditionally done just this, building up animation from sequences of hand mo…

Cited by 69PDFScholar
2021

SeeHear: Signer Diarisation and a New Dataset

ICASSP 2021accepted

In this work, we propose a framework to collect a large-scale, diverse sign language dataset that can be used to train automatic sign language recognition models.The first contribution of this work is SDTrack, a generic method for signer tracking and diarisation in the wild. Our second contribution…

Cited by 0SourceScholar
2021

There and Back Again: Self-supervised Multispectral Correspondence Estimation

ICRA 2021poster

Across a wide range of applications, from autonomous vehicles to medical imaging, multi-spectral images provide an opportunity to extract additional information not present in color images. One of the most important steps in making this information readily available is the accurate estimation of den…

Cited by 11SourceScholar
2021

VDSM: Unsupervised Video Disentanglement With State-Space Modeling and Deep Mixtures of Experts

CVPR 2021poster

Disentangled representations support a range of downstream tasks including causal reasoning, generative modeling, and fair machine learning. Unfortunately, disentanglement has been shown to be impossible without the incorporation of supervision or inductive bias. Given that supervision is often expe…

Cited by 10PDFcodeScholar
2020

DeFeat-Net: General Monocular Depth via Simultaneous Unsupervised Representation Learning

CVPR 2020poster

In the current monocular depth research, the dominant approach is to employ unsupervised training on large datasets, driven by warped photometric consistency. Such approaches lack robustness and are unable to generalize to challenging domains such as nighttime scenes or adverse weather conditions wh…

Cited by 106PDFcodeScholar
2020

Progressive Transformers for End-to-End Sign Language Production

ECCV 2020poster

The goal of automatic Sign Language Production (SLP) is to translate spoken language to a continuous stream of sign language video at a level comparable to a human translator. If this was achievable, then it would revolutionise Deaf hearing communications. Previous work on predominantly isolated SLP…

2020

Same Features, Different Day: Weakly Supervised Feature Learning for Seasonal Invariance

CVPR 2020poster

"Like night and day" is a commonly used expression to imply that two things are completely different. Unfortunately, this tends to be the case for current visual feature representations of the same scene across varying seasons or times of day. The aim of this paper is to provide a dense feature repr…

Cited by 27PDFcodeScholar
2020

Sign Language Transformers: Joint End-to-End Sign Language Recognition and Translation

CVPR 2020oral

Prior work on Sign Language Translation has shown that having a mid-level sign gloss representation (effectively recognizing the individual signs) improves the translation performance drastically. In fact, the current state-of-the-art in translation requires gloss level tokenization in order to work…

Cited by 726PDFcodeScholar
2020

Training Adversarial Agents to Exploit Weaknesses in Deep Control Policies

ICRA 2020poster

Deep learning has become an increasingly common technique for various control problems, such as robotic arm manipulation, robot navigation, and autonomous vehicles. However, the downside of using deep neural networks to learn control policies is their opaque nature and the difficulties of validating…

Cited by 64SourcecodeScholar
2019

A Robust Extrinsic Calibration Framework for Vehicles with Unscaled Sensors

IROS 2019poster

Accurate extrinsic sensor calibration is essential for both autonomous vehicles and robots. Traditionally this is an involved process requiring calibration targets, known fiducial markers and is generally performed in a lab. Moreover, even a small change in the sensor layout requires recalibration.…

Cited by 11SourceScholar
2019

HMM-based Approaches to Model Multichannel Information in Sign Language Inspired from Articulatory Features-based Speech Processing

ICASSP 2019accepted

Sign language conveys information through multiple channels, such as hand shape, hand movement, and mouthing. Modeling this multichannel information is a highly challenging problem. In this paper, we elucidate the link between spoken language and sign language in terms of production phenomenon and p…

Cited by 0SourceScholar
2019

Scale-Adaptive Neural Dense Features: Learning via Hierarchical Context Aggregation

CVPR 2019poster

How do computers and intelligent agents view the world around them? Feature extraction and representation constitutes one the basic building blocks towards answering this question. Traditionally, this has been done with carefully engineered hand-crafted techniques such as HOG, SIFT or ORB. However,…

Cited by 16PDFcodeScholar
2018

Neural Sign Language Translation

CVPR 2018poster

Sign Language Recognition (SLR) has been an active research field for the last two decades. However, most research to date has considered SLR as a naive gesture recognition problem. SLR seeks to recognize a sequence of continuous signs but neglects the underlying rich grammatical and linguistic stru…

2018

SeDAR - Semantic Detection and Ranging: Humans can Localise without LiDAR, can Robots?

ICRA 2018poster

How does a person work out their location using a floorplan? It is probably safe to say that we do not explicitly measure depths to every visible surface and try to match them against different pose estimates in the floorplan. And yet, this is exactly how most robotic scan-matching algorithms operat…

Cited by 47SourceScholar
2017

SubUNets: End-To-End Hand Shape and Continuous Sign Language Recognition

ICCV 2017spotlight

We propose a novel deep learning approach to solve simultaneous alignment and recognition problems (referred to as "Sequence-to-sequence" learning). We decompose the problem into a series of specialised expert systems referred to as SubUNets. The spatio-temporal relationships between these SubUNets…

Cited by 420PDFcodeScholar
2017

Taking the Scenic Route to 3D: Optimising Reconstruction From Moving Cameras

ICCV 2017poster

Reconstruction of 3D environments is a problem that has been widely addressed in the literature. While many approaches exist to perform reconstruction, few of them take an active role in deciding where the next observations should come from. Furthermore, the problem of travelling from the camera's c…

Cited by 25PDFScholar
2016

Deep Hand: How to Train a CNN on 1 Million Hand Images When Your Data Is Continuous and Weakly Labelled

CVPR 2016oral

This work presents a new approach to learning a frame-based classifier on weakly labelled sequence data by embedding a CNN within an iterative EM algorithm. This allows the CNN to be trained on a vast number of example images when only loose sequence level information is available for the source vid…

Cited by 370PDFScholar