← Search

Yucheng Suo

4 accepted papers

2025

From Trial to Triumph: Advancing Long Video Understanding via Visual Context Sample Scaling and Self-reward Alignment

ICCV 2025poster

Multi-modal Large language models (MLLMs) show remarkable ability in video understanding. Nevertheless, understanding long videos remains challenging as the models can only process a finite number of frames in a single inference, potentially omitting crucial visual information. To address the challe…

Cited by 0SourcePDFScholar
2025

Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflection

ICLR 2025poster

Visual instructions for long-horizon tasks are crucial as they intuitively clarify complex concepts and enhance retention across extended steps. Directly generating a series of images using text-to-image models without considering the context of previous steps results in inconsistent images, increa…

Cited by 0SourcePDFScholar
2024

Knowledge-Enhanced Dual-stream Zero-shot Composed Image Retrieval

CVPR 2024poster

We study the zero-shot Composed Image Retrieval (ZS-CIR) task which is to retrieve the target image given a reference image and a description without training on the triplet datasets. Previous works generate pseudo-word tokens by projecting the reference image features to the text embedding space. H…