← Search

Raphael Schumann

4 accepted papers

2024

VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View

AAAI 2024technical

Incremental decision making in real-world environments is one of the most challenging tasks in embodied artificial intelligence. One particularly demanding scenario is Vision and Language Navigation (VLN) which requires visual and natural language understanding as well as spatial and temporal reason…

2022

Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas

ACL 2022long

Vision and language navigation (VLN) is a challenging visually-grounded language understanding task. Given a natural language navigation instruction, a visual agent interacts with a graph-based environment equipped with panorama images and tries to follow the described route. Most prior work has bee…

2018

Incorporating ASR Errors with Attention-Based, Jointly Trained RNN for Intent Detection and Slot Filling

ICASSP 2018accepted

The real-world performance of slot filling and intent detection task generally degrades due to transcription errors generated by speech recognition engine. The insertion, deletion, and mis-recognition errors from speech recognizer's front-end cause the mis-interpretation and mis-alignment of the lan…

Cited by 0SourceScholar