← Search

Zuyao You

3 accepted papers

2026

VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding

ICML 2026poster

Recent advancements in Video Large Language Models (Video LLMs) have demonstrated impressive results, yet existing approaches handle either temporal or spatial dimension in isolation, struggling in the analysis of complex events that require spatial-temporal integration. To bridge this gap, we propo…

Cited by 5SourceScholar
2024

Synthesize Diagnose and Optimize: Towards Fine-Grained Vision-Language Understanding

CVPR 2024poster

Vision language models (VLM) have demonstrated remarkable performance across various downstream tasks. However understanding fine-grained visual-linguistic concepts such as attributes and inter-object relationships remains a significant challenge. While several benchmarks aim to evaluate VLMs in fin…