← Search

Nina Shvetsova*

1 accepted papers

2024

HowToCaption: Prompting LLMs to Transform Video Annotations at Scale

ECCV 2024poster

"Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR) from the audio signal in the videos. However, in contrast to human-annotated captions, both speech and subtitles natu…