← Search

Sathish Reddy Indurthi

10 accepted papers

2025

Cross-lingual Evaluation of Multilingual Text Generation

COLING 2025main

Scaling automatic evaluation of multilingual text generation of LLMs to new tasks, domains, and languages remains a challenge. Traditional evaluation on benchmark datasets carries the risk of reference data leakage in LLM training or involves additional human annotation effort. The alternative strat…

2025

Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy

ICLR 2025poster

Large Language Models (LLMs) are susceptible to security and safety threats, such as prompt injection, prompt extraction, and harmful requests. One major cause of these vulnerabilities is the lack of an instruction hierarchy. Modern LLM architectures treat all inputs equally, failing to distinguish…

Cited by 6SourcePDFScholar
2024

Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets

EMNLP 2024finding

Advancements in Large Language Models (LLMs) have significantly enhanced instruction-following capabilities. However, most Instruction Fine-Tuning (IFT) datasets are predominantly in English, limiting model performance in other languages. Traditional methods for creating multilingual IFT datasets—su…

2024

WPO: Enhancing RLHF with Weighted Preference Optimization

EMNLP 2024main

Reinforcement learning from human feedback (RLHF) is a promising solution to align large language models (LLMs) more closely with human values. Off-policy preference optimization, where the preference data is obtained from other models, is widely adopted due to its cost efficiency and scalability. H…

2023

CLAD-ST: Contrastive Learning with Adversarial Data for Robust Speech Translation

EMNLP 2023short main

The cascaded approach continues to be the most popular choice for speech translation (ST). This approach consists of an automatic speech recognition (ASR) model and a machine translation (MT) model that are used in a pipeline to translate speech in one language to text in another language. MT models…

Cited by 0SourceScholar
2023

Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations

EMNLP 2023short main

Large language models (LLMs) like ChatGPT can be expensive to train, deploy, and use for specific natural language generation tasks such as text summarization and for certain domains. A promising alternative is to fine-tune relatively smaller language models (LMs) on a particular task using high-qua…

Cited by 0SourceScholar
2022

Language Model Augmented Monotonic Attention for Simultaneous Translation

NAACL 2022long

The state-of-the-art adaptive policies for Simultaneous Neural Machine Translation (SNMT) use monotonic attention to perform read/write decisions based on the partial source and target sequences. The lack of sufficient information might cause the monotonic attention to take poor read/write decisions…

Cited by 9SourcePDFScholar
2021

Task Aware Multi-Task Learning for Speech to Text Tasks

ICASSP 2021accepted

In general, the direct Speech-to-text translation (ST) is jointly trained with Automatic Speech Recognition (ASR), and Machine Translation (MT) tasks. However, the issues with the current joint learning strategies inhibit the knowledge transfer across these tasks. We propose a task modulation networ…

Cited by 0SourceScholar
2020

End-end Speech-to-Text Translation with Modality Agnostic Meta-Learning

ICASSP 2020accepted

Collecting large amounts of data to train end-to-end Speech Translation (ST) models is more difficult compared to the ASR and MT tasks. Previous studies have proposed the use of transfer learning approaches to overcome the above difficulty. These approaches benefit from weakly supervised training da…

Cited by 0SourceScholar
2020

Small Energy Masking for Improved Neural Network Training for End-To-End Speech Recognition

ICASSP 2020accepted

In this paper, we present a Small Energy Masking (SEM) algorithm, which masks inputs having values below a certain threshold. More specifically, a time-frequency bin is masked if the filterbank energy in this bin is less than a certain energy threshold. A uniform distribution is employed to randomly…

Cited by 0SourceScholar