← Search

Jinsong Lan

8 accepted papers

2026

Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement Learning

CVPR 2026

Large Vision-Language Models (LVLMs) often omit or misrepresent critical visual content in generated image captions. Minimizing such information loss will force LVLMs to focus on image details to generate precise descriptions. However, measuring information loss during modality conversion is inheren

Cited by 0SourceScholar
2026

ORION: Decoupling and Alignment for Unified Autoregressive Understanding and Generation

ICLR 2026poster

Unified multimodal Large Language Models (MLLMs) hold great promise for seamlessly integrating understanding and generation. However, monolithic autoregressive architectures, despite their elegance and conversational fluency, suffer from a fundamental semantic–structural conflict: optimizing for low…

Cited by 0SourceScholar
2026

iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance

ICML 2026poster

Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal consistency, they are predominantly confined to non-interactive scenarios where models merely showcase garments. This li…

Cited by 0SourceScholar
2025

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training

CVPR 2025poster

In rapidly evolving field of vision-language models (VLMs), contrastive language-image pre-training (CLIP) has made significant strides, becoming foundation for various downstream tasks. However, relying on one-to-one (image, text) contrastive paradigm to learn alignment from large-scale messy web d…

2025

INTER: Mitigating Hallucination in Large Vision-Language Models by Interaction Guidance Sampling

ICCV 2025poster

Hallucinations in large vision-language models (LVLMs) pose significant challenges for real-world applications, as LVLMs may generate responses that appear plausible yet remain inconsistent with the associated visual content. This issue rarely occurs in human cognition. We argue that this discrepanc…

2025

Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal Model

AAAI 2025technical

Large Multimodal Models (LMMs) have significantly progressed by extending large language models. Building on this progress, the latest developments in LMMs demonstrate the ability to generate dense pixel-wise segmentation by integrating segmentation models. Despite the innovations, existing works’ t…

2024

Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models

ECCV 2024oral

"Vision-Language Large Models (VLMs) recently become primary backbone of AI, due to the impressive performance. However, their expensive computation costs, i.e., throughput and delay, impede potentials in the real-world scenarios. To achieve acceleration for VLMs, most existing methods focus on the…

Cited by 9SourcePDFScholar
2024

Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment

ECCV 2024poster

"This paper introduces a novel framework for virtual try-on, termed . Different from previous methods, is a customizable solution. Besides generating high-fidelity results, our method supports users to precisely manipulate the wearing style. To achieve this goal, we first construct a strong pipeline…