← Search

Jianting Tang

2 accepted papers

2025

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

ICCV 2025poster

Mainstream Multimodal Large Language Models (MLLMs) achieve visual understanding by using a vision projector to bridge well-pretrained vision encoders and large language models (LLMs). The inherent gap between visual and textual modalities makes the embeddings from the vision projector critical for…

Cited by 0SourcePDFScholar
2025

Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images

ICLR 2025poster

Large vision-language models (LVLMs) have demonstrated remarkable image understanding and dialogue capabilities, allowing them to handle a variety of visual question answering tasks. However, their widespread availability raises concerns about unauthorized usage and copyright infringement, where use…

Cited by 1SourcePDFScholar