MeToM: Metadata-Guided Token Merging for Efficient Video LLMs
Video Large Language Models (VLLMs) encounter significant computational challenges due to the large volume of visual tokens generated from multiple frames. Existing visual token pruning methods fail to account for the uneven spatiotemporal information density, thus squandering scarce token budgets o