← Search

Vishruit Kulshreshtha

1 accepted papers

2025

VIT-Pro: Visual Instruction Tuning for Product Images

NAACL 2025industry

General vision-language models (VLMs) trained on web data struggle to understand and converse about real-world e-commerce product images. We propose a cost-efficient approach for collecting training data to train a generative VLM for e-commerce product images. The key idea is to leverage large-scale…

Cited by 0SourcePDFScholar