2024
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
ECCV 2024poster
"Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR) from the audio signal in the videos. However, in contrast to human-annotated captions, both speech and subtitles natu…