← Search

Yubo Zhu

3 accepted papers

2026

UVU: Improving Multimodal Understanding via Vision-Language Unified Autoregressive Paradigm

CVPR 2026

Despite remarkable advancements in multimodal large language models (MLLMs), their fine-grained visual understanding is constrained by a primary reliance on sparse textual supervision. Existing efforts to introduce visual supervision typically do so during post-training, when visual representations

Cited by 0SourceScholar
2025

The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations

EMNLP 2025

Estimating the difficulty of input questions as perceived by large language models (LLMs) is essential for accurate performance evaluation and adaptive inference. Existing methods typically rely on repeated response sampling, auxiliary models, or fine-tuning the target model itself, which may incur

Cited by 0SourcePDFScholar
2025

Video Summarization Using Denoising Diffusion Probabilistic Model

AAAI 2025technical

Video summarization aims to eliminate visual redundancy while retaining key parts of video to construct concise and comprehensive synopses. Most existing methods use discriminative models to predict the importance scores of video frames. However, these methods are susceptible to annotation inconsist…

Cited by 0SourcePDFScholar