← Search

Yiyu Wang

6 accepted papers

2026

Accelerating Streaming Video Large Language Models via Hierarchical Token Compression

CVPR 2026

Streaming Video Large Language Models (VideoLLMs) have demonstrated impressive performance across various video understanding tasks, but they face significant challenges in real-time deployment due to the high computational cost of processing dense visual tokens from continuous video streams. In str

Cited by 0SourcecodeScholar
2026

Variation-aware Vision Token Dropping for Faster Large Vision-Language Models

CVPR 2026

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding tasks. However, the increasing demand for high-resolution image and long-video understanding results in substantial token counts, consequently leading to reduced inference efficiency. Token com

Cited by 0SourcecodeScholar
2025

A Unified Agentic Framework for Evaluating Conditional Image Generation

ACL 2025long

Conditional image generation has gained significant attention for its ability to personalize content. However, the field faces challenges in developing task-agnostic, reliable, and explainable evaluation metrics. This paper introduces CIGEval, a unified agentic framework for comprehensive evaluation…

2025

Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models

EMNLP 2025

Video large language models (VideoLLM) excel at video understanding, but face efficiency challenges due to the quadratic complexity of abundant visual tokens. Our systematic analysis of token compression methods for VideoLLMs reveals two critical issues: (i) overlooking distinctive visual signals ac

2020

Neural Topographic Factor Analysis for fMRI Data

NeurIPS 2020poster

Neuroimaging studies produce gigabytes of spatio-temporal data for a small number of participants and stimuli. Recent work increasingly suggests that the common practice of averaging across participants and stimuli leaves out systematic and meaningful information. We propose Neural Topographic Facto…

Cited by 9SourcePDFScholar