← Search

Jamie Shotton

19 accepted papers

2024

Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving

ICRA 2024poster

Large Language Models (LLMs) have shown promise in the autonomous driving sector, particularly in generalization and interpretability. We introduce a unique objectlevel multimodal LLM architecture that merges vectorized numeric modalities with a pre-trained LLM to improve context understanding in dr…

Cited by 229SourcecodeScholar
2024

LingoQA: Video Question Answering for Autonomous Driving

ECCV 2024poster

"We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capab…

2022

Model-Based Imitation Learning for Urban Driving

NeurIPS 2022accept

An accurate model of the environment and the dynamic agents acting in it offers great potential for improving motion planning. We present MILE: a Model-based Imitation LEarning approach to jointly learn a model of the world and a policy for autonomous driving. Our method leverages 3D geometry as an…

2021

Fake It Till You Make It: Face Analysis in the Wild Using Synthetic Data Alone

ICCV 2021poster

We demonstrate that it is possible to perform face-related computer vision in the wild using synthetic data alone. The community has long enjoyed the benefits of synthesizing training data with graphics, but the domain gap between real and synthetic data has remained a problem, especially for human…

Cited by 339PDFcodeScholar
2021

FastNeRF: High-Fidelity Neural Rendering at 200FPS

ICCV 2021poster

Recent work on Neural Radiance Fields (NeRF) showed how neural networks can be used to encode complex 3D environments that can be rendered photorealistically from novel viewpoints. Rendering these images is very computationally demanding and recent improvements are still a long way from enabling int…

Cited by 787PDFScholar
2021

Full-Body Motion From a Single Head-Mounted Device: Generating SMPL Poses From Partial Observations

ICCV 2021poster

The increased availability and maturity of head-mounted and wearable devices opens up opportunities for remote communication and collaboration. However, the signal streams provided by these devices (e.g., head pose, hand pose, and gaze direction) do not represent a whole person. One of the main open…

Cited by 69PDFScholar
2020

CONFIG: Controllable Neural Face Image Generation

ECCV 2020poster

Our ability to sample realistic natural images, particularly faces, has advanced by leaps and bounds in recent years, yet our ability to exert fine-tuned control over the generative process has lagged behind. If this new technology is to find practical uses, we need to achieve a level of control ove…

2020

High Resolution Zero-Shot Domain Adaptation of Synthetically Rendered Face Images

ECCV 2020poster

Generating photorealistic images of human faces at scale remains a prohibitively difficult task using computer graphics approaches. This is because these require the simulation of light to be photorealistic, which in turn requires physically accurate modelling of geometry, materials, and light sourc…

Cited by 11SourcePDFScholar
2020

The Phong Surface: Efficient 3D Model Fitting using Lifted Optimization

ECCV 2020poster

Realtime perceptual and interaction capabilities in mixed reality require a range of 3D tracking problems to be solved at low latency on resource-constrained hardware such as head-mounted devices. Indeed, for devices such as HoloLens 2 where the CPU and GPU are left available for applications, multi…

Cited by 17SourcePDFScholar
2019

Learning-driven Coarse-to-Fine Articulated Robot Tracking

ICRA 2019poster

In this work we present an articulated tracking approach for robotic manipulators, which relies only on visual cues from colour and depth images to estimate the robot's state when interacting with or being occluded by its environment. We hypothesise that articulated model fitting approaches can only…

Cited by 8SourceScholar
2018

Visual Articulated Tracking in the Presence of Occlusions

ICRA 2018poster

This paper focuses on visual tracking of a robotic manipulator during manipulation. In this situation, tracking is prone to failure when visual distractions are created by the object being manipulated and the clutter in the environment. Current state-of-the-art approaches, which typically rely on mo…

Cited by 7SourceScholar
2017

DSAC - Differentiable RANSAC for Camera Localization

CVPR 2017oral

RANSAC is an important algorithm in robust optimization and a central building block for many computer vision applications. In recent years, traditionally hand-crafted pipelines have been replaced by deep learning pipelines, which can be trained in an end-to-end fashion. However, RANSAC has so far n…

Cited by 737PDFcodeScholar
2017

PoseAgent: Budget-Constrained 6D Object Pose Estimation via Reinforcement Learning

CVPR 2017poster

State-of-the-art computer vision algorithms often achieve efficiency by making discrete choices about which hypotheses to explore next. This allows allocation of computational resources to promising candidates, however, such decisions are non-differentiable. As a result, these algorithms are hard t…

Cited by 60PDFScholar
2016

Fits Like a Glove: Rapid and Reliable Hand Shape Personalization

CVPR 2016spotlight

We present a fast, practical method for personalizing a hand shape basis to an individual user's detailed hand shape using only a small set of depth images. To achieve this, we minimize an energy based on a sum of render-and-compare cost functions called the golden energy. However, this energy is on…

Cited by 154PDFScholar
2015

Depth-Based Hand Pose Estimation: Data, Methods, and Challenges

ICCV 2015poster

Hand pose estimation has matured rapidly in recent years. The introduction of commodity depth sensors and a multitude of practical applications have spurred new advances. We provide an extensive analysis of the state-of-the-art, focusing on hand pose estimation from a single depth frame. To do so, w…

Cited by 200PDFScholar
2015

Exploiting Uncertainty in Regression Forests for Accurate Camera Relocalization

CVPR 2015poster

Recent advances in camera relocalization use predictions from a regression forest to guide the camera pose optimization procedure. In these methods, each tree associates one pixel with a point in the scene's 3D world coordinate frame. In previous work, these predictions were point estimates and the…

Cited by 192SourcePDFScholar
2015

Learning an Efficient Model of Hand Shape Variation From Depth Images

CVPR 2015poster

We describe how to learn a compact and efficient model of the surface deformation of human hands. The model is built from a set of noisy depth images of a diverse set of subjects performing different poses with their hands. We represent the observed surface using Loop subdivision of a control mesh t…

Cited by 163SourcePDFScholar
2015

Model-Based Tracking at 300Hz Using Raw Time-of-Flight Observations

ICCV 2015poster

Consumer depth cameras have dramatically improved our ability to track rigid, articulated, and deformable 3D objects in real-time. However, depth cameras have a limited temporal resolution (frame-rate) that restricts the accuracy and robustness of tracking, especially for fast or unpredictable motio…

Cited by 23PDFScholar
2015

Opening the Black Box: Hierarchical Sampling Optimization for Estimating Human Hand Pose

ICCV 2015oral

We address the problem of hand pose estimation, formulated as an inverse problem. Typical approaches optimize an energy function over pose parameters using a `black box' image generation procedure. This procedure knows little about either the relationships between the parameters or the form of the…

Cited by 170PDFScholar