2024
LITA: Language Instructed Temporal-Localization Assistant
ECCV 2024poster
"There has been tremendous progress in multimodal Large Language Models (LLMs). Recent works have extended these models to video input with promising instruction following capabilities. However, an important missing piece is temporal localization. These models cannot accurately answer the “When?” qu…