← Search

Xudong Hong

3 accepted papers

2024

HowToCaption: Prompting LLMs to Transform Video Annotations at Scale

ECCV 2024poster

"Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR) from the audio signal in the videos. However, in contrast to human-annotated captions, both speech and subtitles natu…

2024

Retrieval-Augmented Modular Prompt Tuning for Low-Resource Data-to-Text Generation

COLING 2024main

Data-to-text (D2T) generation describes the task of verbalizing data, often given as attribute-value pairs. While this task is relevant for many different data domains beyond the traditionally well-explored tasks of weather forecasting, restaurant recommendations, and sports reporting, a major chall…

2023

Visual Coherence Loss for Coherent and Visually Grounded Story Generation

ACL 2023findings

Local coherence is essential for long-form text generation models. We identify two important aspects of local coherence within the visual storytelling task: (1) the model needs to represent re-occurrences of characters within the image sequence in order to mention them correctly in the story; (2) ch…