← Search

Linbo Jin

5 accepted papers

2025

ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning

ACL 2025finding

Direct Preference Optimization (DPO) has gained significant attention for its simplicity and computational efficiency in aligning large language models (LLMs). Recent advancements have extended DPO to multimodal scenarios, achieving strong performance. However, traditional DPO relies on binary prefe…

Cited by 0SourcePDFScholar
2025

Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples

AAAI 2025technical

Existing Vision-Language Pretraining (VLP) methods have achieved remarkable improvements across a variety of vision-language tasks, confirming their effectiveness in capturing coarse-grained semantic correlations. However, their capability for fine-grained understanding, which is critical for many…

Cited by 2SourcePDFScholar
2024

MoDULA: Mixture of Domain-Specific and Universal LoRA for Multi-Task Learning

EMNLP 2024main

The growing demand for larger-scale models in the development of Large Language Models (LLMs) poses challenges for efficient training within limited computational resources. Traditional fine-tuning methods often exhibit instability in multi-task learning and rely heavily on extensive training resour…

Cited by 1SourcePDFScholar
2023

FashionKLIP: Enhancing E-Commerce Image-Text Retrieval with Fashion Multi-Modal Conceptual Knowledge Graph

ACL 2023industry

Image-text retrieval is a core task in the multi-modal domain, which arises a lot of attention from both research and industry communities. Recently, the booming of visual-language pre-trained (VLP) models has greatly enhanced the performance of cross-modal retrieval. However, the fine-grained inter…

2021

Kaleido-BERT: Vision-Language Pre-Training on Fashion Domain

CVPR 2021poster

We present a new vision-language (VL) pre-training model dubbed Kaleido-BERT, which introduces a novel kaleido strategy for fashion cross-modality representations from transformers. In contrast to random masking strategy of recent VL models, we design alignment guided masking to jointly focus more o…

Cited by 154PDFcodeScholar