← Search

Shamil Chollampatt

4 accepted papers

2025

Cross-lingual Evaluation of Multilingual Text Generation

COLING 2025main

Scaling automatic evaluation of multilingual text generation of LLMs to new tasks, domains, and languages remains a challenge. Traditional evaluation on benchmark datasets carries the risk of reference data leakage in LLM training or involves additional human annotation effort. The alternative strat…

2024

Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets

EMNLP 2024finding

Advancements in Large Language Models (LLMs) have significantly enhanced instruction-following capabilities. However, most Instruction Fine-Tuning (IFT) datasets are predominantly in English, limiting model performance in other languages. Traditional methods for creating multilingual IFT datasets—su…

2023

CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation

EMNLP 2023short main

The cascaded approach continues to be the most popular choice for speech translation (ST). This approach consists of an automatic speech recognition (ASR) model and a machine translation (MT) model that are used in a pipeline to translate speech in one language to text in another language. MT models…

Cited by 0SourceScholar
2023

Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations

EMNLP 2023short main

Large language models (LLMs) like ChatGPT can be expensive to train, deploy, and use for specific natural language generation tasks such as text summarization and for certain domains. A promising alternative is to fine-tune relatively smaller language models (LMs) on a particular task using high-qua…

Cited by 0SourceScholar