← Search

Vijay Badrinarayanan

8 accepted papers

2024

LingoQA: Video Question Answering for Autonomous Driving

ECCV 2024poster

"We introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving. The dataset contains 28K unique short video scenarios, and 419K annotations. Evaluating state-of-the-art vision-language models on our benchmark shows that their performance is below human capab…

2021

FIERY: Future Instance Prediction in Bird's-Eye View From Surround Monocular Cameras

ICCV 2021poster

Driving requires interacting with road agents and predicting their future behaviour in order to navigate safely. We present FIERY: a probabilistic future prediction model in bird's-eye view from monocular cameras. Our model predicts future instance segmentation and motion of dynamic agents that can…

Cited by 330PDFcodeScholar
2020

Atlas: End-to-End 3D Scene Reconstruction from Posed Images

ECCV 2020poster

We present an end-to-end 3D reconstruction of a scene by directly regressing a truncated signed distance function (TSDF) from a set of posed RGB images. Traditional approaches to 3D reconstruction rely on an intermediate representation of depth maps prior to estimating a full 3D model of a scene. We…

2020

DELTAS: Depth Estimation by Learning Triangulation And densification of Sparse points

ECCV 2020poster

Multi-view stereo (MVS) is the golden mean between the accuracy of active depth sensing and the practicality of monocular depth estimation. Cost volume based approaches employing 3D convolutional neural networks (CNNs) have considerably improved the accuracy of MVS systems. However, this accuracy co…

2018

GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks

ICML 2018oral

Deep multitask networks, in which one neural network produces multiple predictive outputs, can offer better speed and performance than their single-task counterparts but are challenging to train properly. We present a gradient normalization (GradNorm) algorithm that automatically balances training i…

Cited by 1623SourcePDFScholar
2016

Understanding Real World Indoor Scenes With Synthetic Data

CVPR 2016poster

Scene understanding is a prerequisite to many high level tasks for any automated intelligent machine operating in real world environments. Recent attempts with supervised learning have shown promise in this direction but also highlighted the need for enormous quantity of supervised data --- performa…

Cited by 448PDFScholar