GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
We focus on the task of identifying the location of target regions from a natural language instruction and a front camera image captured by a mobility. This task is challenging because it requires both existence prediction and segmentation mask generation, particularly for stuff-type target regions…