EMNLP 2022main4 citations

ALFRED-L: Investigating the Role of Language for Action Learning in Interactive Visual Environments

Arjun Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tur

Abstract

Embodied Vision and Language Task Completion requires an embodied agent to interpret natural language instructions and egocentric visual observations to navigate through and interact with environments. In this work, we examine ALFRED, a challenging benchmark for embodied task completion, with the goal of gaining insight into how effectively models utilize language. We find evidence that sequence-to-sequence and transformer-based models trained on this benchmark are not sufficiently sensitive to changes in input language instructions. Next, we construct a new test split – ALFRED-L to test whether ALFRED models can generalize to task structures not seen during training that intuitively require the same types of language understanding required in ALFRED. Evaluation of existing models on ALFRED-L suggests that (a) models are overly reliant on the sequence in which objects are visited in typical ALFRED trajectories and fail to adapt to modifications of this sequence and (b) models trained with additional augmented trajectories are able to adapt relatively better to such changes in input language instructions.

BibTeX
@inproceedings{akula-etal-2022-alfred,
    title = "{ALFRED}-{L}: Investigating the Role of Language for Action Learning in Interactive Visual Environments",
    author = "Akula, Arjun  and
      Gella, Spandana  and
      Padmakumar, Aishwarya  and
      Namazifar, Mahdi  and
      Bansal, Mohit  and
      Thomason, Jesse  and
      Hakkani-Tur, Dilek",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.636/",
    doi = "10.18653/v1/2022.emnlp-main.636",
    pages = "9369--9378"
}
ALFRED-L: Investigating the Role of Language for Action Learning in Interactive Visual Environments · EMNLP 2022