NeuralFloors: Conditional Street-Level Scene Generation From BEV Semantic Maps via Neural Fields
Valentina Musat, Daniele De Martini, Matthew Gadd, Paul Newman
Abstract
Semantic Bird's Eye View (BEV) representations are a popular format, being easily interpretable and editable. However, synthesising ground-view images from BEVs is a difficult task as the system would need to learn both the mapping from BEV to Front View (FV) structure as well as to synthesise highly photo-realistic imagery, thus having to simultaneously consider both the geometry and appearance of the scene. We therefore present a factorised approach that tackles the problem in two stages: a first stage that learns a BEV to FV transformation in the semantic space through a Neural Field, and a second stage that leverages a Latent Diffusion Model (LDM) to synthesise images conditional on the output of the first stage. Our experiments show that this approach produces RGB images with a high perceptual quality that are also well aligned with their corresponding FV ground-truth.
BibTeX
@inproceedings{ral2024_neuralfloorscond,
title = {NeuralFloors: Conditional Street-Level Scene Generation From BEV Semantic Maps via Neural Fields},
author = {Valentina Musat and Daniele De Martini and Matthew Gadd and Paul Newman},
booktitle = {RA-L 2024},
year = {2024}
}