← Search

Alex Kendall

13 accepted papers

2024

LingoQA: Video Question Answering for Autonomous Driving

ECCV 2024poster

"We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capab…

2022

Model-Based Imitation Learning for Urban Driving

NeurIPS 2022accept

An accurate model of the environment and the dynamic agents acting in it offers great potential for improving motion planning. We present MILE: a Model-based Imitation LEarning approach to jointly learn a model of the world and a policy for autonomous driving. Our method leverages 3D geometry as an…

2021

FIERY: Future Instance Prediction in Bird's-Eye View From Surround Monocular Cameras

ICCV 2021poster

Driving requires interacting with road agents and predicting their future behaviour in order to navigate safely. We present FIERY: a probabilistic future prediction model in bird's-eye view from monocular cameras. Our model predicts future instance segmentation and motion of dynamic agents that can…

Cited by 330PDFcodeScholar
2020

Probabilistic Future Prediction for Video Scene Understanding

ECCV 2020poster

We present a novel deep learning architecture for probabilistic future prediction from video. We predict the future semantics, geometry and motion of complex real-world urban scenes and use this representation to control an autonomous vehicle. This work is the first to jointly predict ego-motion, st…

Cited by 95SourcePDFScholar
2019

Learning to Drive from Simulation without Real World Labels

ICRA 2019poster

Simulation can be a powerful tool for under-standing machine learning systems and designing methods to solve real-world problems. Training and evaluating methods purely in simulation is often “doomed to succeed” at the desired task in a simulated environment, but the resulting models are incapable o…

Cited by 146SourceScholar
2019

Learning to Drive in a Day

ICRA 2019poster

We demonstrate the first application of deep reinforcement learning to autonomous driving. From randomly initialised parameters, our model is able to learn a policy for lane following in a handful of training episodes using a single monocular image as input. We provide a general and easy to obtain r…

Cited by 956SourceScholar
2018

Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics

CVPR 2018poster

Numerous deep learning applications benefit from multi-task learning with multiple regression and classification objectives. In this paper we make the observation that the performance of such systems is strongly dependent on the relative weighting between each task's loss. Tuning these weights by ha…

Cited by 4160SourcePDFScholar
2017

End-To-End Learning of Geometry and Context for Deep Stereo Regression

ICCV 2017spotlight

We propose a novel deep learning architecture for regressing disparity from a rectified pair of stereo images. We leverage knowledge of the problem's geometry to form a cost volume using deep feature representations. We learn to incorporate contextual information using 3-D convolutions over this vol…

Cited by 1619PDFScholar
2015

PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization

ICCV 2015poster

We present a robust and real-time monocular six degree of freedom relocalization system. Our system trains a convolutional neural network to regress the 6-DOF camera pose from a single RGB image in an end-to-end manner with no need of additional engineering or graph optimisation. The algorithm can o…

Cited by 2991PDFScholar