← Search

Yumeng Shi

5 accepted papers

2026

Causality Matters: How Temporal Information Emerges in Video Language Models

AAAI 2026technical

Video language models (VideoLMs) have made significant progress in multimodal understanding. However, temporal understanding, which involves identifying event order, duration, and relationships across time, still remains a core challenge. Prior works emphasize positional encodings (PEs) as a key mec

Cited by 0SourcePDFScholar
2026

LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play Toolkit

AAAI 2026technical

Large Vision-Language Models (VLMs) exhibit impressive multi-modal capabilities but suffer from prohibitive computational and memory demands, due to their long visual token sequences and massive parameter sizes. To address these issues, recent works have proposed training-free compression methods. H

Cited by 0SourcePDFScholar
2026

Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention

ICML 2026poster

Advanced autoregressive (AR) video generation models have improved visual fidelity and interactivity, but the quadratic complexity of attention remains a primary bottleneck for efficient deployment. While existing sparse attention solutions have shown promise on bidirectional models, we identify tha…

Cited by 0SourceScholar
2025

Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering

EMNLP 2025

Video question answering benefits from the rich information in videos, enabling various applications. However, the large volume of tokens generated from long videos presents challenges to memory efficiency and model performance. To alleviate this, existing works propose to compress video inputs, but

2023

Lossy and Lossless (L2) Post-training Model Size Compression

ICCV 2023poster

Deep neural networks have delivered remarkable performance and have been widely used in various visual tasks. However, their huge sizes cause significant inconvenience for transmission and storage. Many previous studies have explored model size compression. However, these studies often approach vari…

Cited by 6PDFcodeScholar