The Visual Prism: Refracting Images into Parallel Multilingual Descriptions with Structured Visual Guidance
Parallel corpora, as the foundation of machine translation, remain crucial even in the era of large language models (LLMs) for pre-training and fine-tuning. However, annotating parallel corpora is extremely costly, as it requires annotators to be proficient in multiple languages. To reduce this cost