← Search

Yan Shu

12 accepted papers

2026

SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion

ICML 2026poster

Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in high-quality image generation. However, achieving fast and accurate inversion—transforming images back to latent noise for faithful reconstruction and editing—remains a challenging bottleneck…

Cited by 0SourceScholar
2026

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

CVPR 2026

Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level visual representations. To address this problem, we introduce TerraScope, a unified VLM that delivers pixel-grounded geospa

Cited by 0SourceScholar
2025

MLVU: Benchmarking Multi-task Long Video Understanding

CVPR 2025poster

The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in vi…

2025

Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

CVPR 2025poster

Long video understanding poses a significant challenge for current Multi-modal Large Language Models (MLLMs). Notably, the MLLMs are constrained by their limited context lengths and the substantial costs while processing long videos. Although several existing methods attempt to reduce visual tokens,…

2025

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

NeurIPS 2025poster

Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visual…

Cited by 0SourceScholar
2024

TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance Control

NeurIPS 2024spotlight

Centred on content modification and style preservation, Scene Text Editing (STE) remains a challenging task despite considerable progress in text-to-image synthesis and text-driven image manipulation recently. GAN-based STE methods generally encounter a common issue of model generalization, while Di…

2024

ZE-FESG: A Zero-Shot Feature Extraction Method Based on Semantic Guidance for No-Reference Video Quality Assessment

ICASSP 2024accepted

Although the current deep neural network based no-reference video quality assessment (NR-VQA) methods can effectively simulate the human visual system (HVS), their interpretability is getting worse. The current methods only extract the low-level features of space and time of the video and do not con…

Cited by 0SourceScholar
2023

Aprogressive Image Dehazing Framework with inter and Intra Contrastive Learning

ICASSP 2023accepted

Image dehazing, aims to estimate latent haze-free images from hazy images, suffering from a lot of lost information. Existing contrastive learning methods tend to utilize hazefree images as positive samples without consideration of negative samples. Even if negative samples are employed, the connect…

Cited by 0SourceScholar
2023

EI2SR: Learning an Enhanced Intra-Instance Semantic Relationship for Arbitrary-Shaped Scene Text Detection

ICASSP 2023accepted

Text detection in natural scenarios, has made significant progress with the deep learning architecture. Towards arbitrary-shaped text detection, fracture detection is the major concern due to the lack of semantic relationship within an instance in existing methods. To circumvent this dilemma, we pro…

Cited by 0SourceScholar
2021

Condensing a Sequence to One Informative Frame for Video Recognition

ICCV 2021poster

Video is complex due to large variations in motion and rich content in fine-grained visual details. Abstracting useful information from such information-intensive media requires exhaustive computing resources. This paper studies a two-step alternative that first condenses the video sequence to an in…

Cited by 10PDFScholar
2020

A Characterization of Mean Squared Error for Estimator with Bagging

AISTATS 2020poster

Bagging can significantly improve the generalization performance of unstable machine learning algorithms such as trees or neural networks. Though bagging is now widely used in practice and many empirical studies have explored its behavior, we still know little about the theoretical properties of bag…