Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive Decoding
A major distinction between video and image understanding is that the former requires reasoning over time. Existing Video Large Language Models (VLLMs) demonstrate promising performance in general video understanding, such as brief captioning or object recognition within individual frames. However,…