IPFormer: Instance Prompt-guided Transformer for Multi-modal Multi-shot Video Understanding
Video Large Language Models (VideoLLMs), which adopt large language models for video understanding, have been demonstrated for single-shot videos. However, they usually struggle in multi-shot videos with frequent shot changes, varying camera angles, etc., which makes VideoLLMs hardly answer question