Visual Grounding Via Heterogeneous Representation Learning and Hierarchical Reasoning of Human-To-Vehicle Commands
With the proliferation of autonomous vehicles (AVs) and their increasing interaction and communication with the riders, how to ground or locate the visual objects of interests (OoIs), such as the concerned pedestrians and other traffic participants, based on the human riders’ natural language and co…