← Search

Shixing Chen

6 accepted papers

2025

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding

NeurIPS 2025poster

Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames. With each video frame typically expanding into tens or hundreds of tokens, the limited context length of large language models (LLMs) forces the VLMs to perceive the frames…

Cited by 0SourceScholar
2023

Movies2Scenes: Using Movie Metadata To Learn Scene Representation

CVPR 2023poster

Understanding scenes in movies is crucial for a variety of applications such as video moderation, search, and recommendation. However, labeling individual scenes is a time-consuming process. In contrast, movie level metadata (e.g., genre, synopsis, etc.) regularly gets produced as part of the film p…

Cited by 19SourcePDFScholar
2021

Shot Contrastive Self-Supervised Learning for Scene Boundary Detection

CVPR 2021poster

Scenes play a crucial role in breaking the storyline of movies and TV episodes into semantically cohesive parts. However, given their complex temporal structure, finding scene boundaries can be a challenging task requiring large amounts of labeled training data. To address this challenge, we present…

Cited by 87PDFScholar