CAVEN: An Embodied Conversational Agent for Efficient Audio-Visual Navigation in Noisy Environments
Audio-visual navigation of an agent towards locating an audio goal is a challenging task especially when the audio is sporadic or the environment is noisy. In this paper, we present CAVEN, a Conversation-based Audio-Visual Embodied Navigation framework in which the agent may interact with a human/o…