2025
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
ICCV 2025poster
Our objective is the automatic generation of Audio Descriptions (ADs) for edited video material, such as movies and TV series. To achieve this, we propose a two-stage framework that leverages "shots" as the fundamental units of video understanding. This includes extending temporal context to neighbo…