← Search

Yingmin Luo

3 accepted papers

2025

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

NeurIPS 2025poster

Recent text-to-image systems face limitations in handling multimodal inputs and complex reasoning tasks. We introduce MindOmni, a unified multimodal large language model that addresses these challenges by incorporating reasoning generation through reinforcement learning. MindOmni leverages a three-p…

Cited by 0SourcecodeScholar
2024

EA-VTR: Event-Aware Video-Text Retrieval

ECCV 2024poster

"Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and the widely adopted video-level cross-modal contrastive learning also struggles to…

Cited by 3SourcePDFScholar
2023

Order-Prompted Tag Sequence Generation for Video Tagging

ICCV 2023poster

Video Tagging intends to infer multiple tags spanning relevant content for a given video. Typically, video tags are freely defined and uploaded by a variety of users, so they have two characteristics: abundant in quantity and disordered intra-video. It is difficult for the existing multi-label class…

Cited by 4PDFScholar