← Search

Chengpeng Fu

5 accepted papers

2026

The Visual Prism: Refracting Images into Parallel Multilingual Descriptions with Structured Visual Guidance

AAAI 2026technical

Parallel corpora, as the foundation of machine translation, remain crucial even in the era of large language models (LLMs) for pre-training and fine-tuning. However, annotating parallel corpora is extremely costly, as it requires annotators to be proficient in multiple languages. To reduce this cost

Cited by 0SourcePDFScholar
2025

Enhancing Non-English Capabilities of English-Centric Large Language Models Through Deep Supervision Fine-Tuning

AAAI 2025technical

Large language models (LLMs) have demonstrated significant progress in multilingual language understanding and generation. However, due to the imbalance in training data, their capabilities in non-English languages are limited. Recent studies revealed the English-pivot multilingual mechanism of LLMs…

2024

Aligning Translation-Specific Understanding to General Understanding in Large Language Models

EMNLP 2024main

Large Language models (LLMs) have exhibited remarkable abilities in understanding complex texts, offering a promising path towards human-like translation performance. However, this study reveals the misalignment between the translation-specific understanding and the general understanding inside LLMs…

2024

Gradient Consistency-based Parameter Allocation for Multilingual Neural Machine Translation

COLING 2024main

Multilingual neural machine translation handles the translation of multiple languages with one unified model. However, this joint-training paradigm incurs the notorious issue of parameter interference, where the model compromises with the language diversity to find a common solution. Recent research…

2023

Enabling Unsupervised Neural Machine Translation with Word-level Visual Representations

EMNLP 2023long findings

Unsupervised neural machine translation has recently made remarkable strides, achieving impressive results with the exclusive use of monolingual corpora. Nonetheless, these methods still exhibit fundamental flaws, such as confusing similar words. A straightforward remedy to rectify this drawback is…

Cited by 0SourceScholar