← Search

Diogo Glória-Silva

2 accepted papers

2024

Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks

ACL 2024long

Multistep instructions, such as recipes and how-to guides, greatly benefit from visual aids, such as a series of images that accompany the instruction steps. While Large Language Models (LLMs) have become adept at generating coherent textual steps, Large Vision/Language Models (LVLMs) are less capab…

Cited by 4SourcePDFScholar
2024

Show and Guide: Instructional-Plan Grounded Vision and Language Model

EMNLP 2024main

Guiding users through complex procedural plans is an inherently multimodal task in which having visually illustrated plan steps is crucial to deliver an effective plan guidance. However, existing works on plan-following language models (LMs) often are not capable of multimodal input and output. In t…