← Search

Abhinav Sethy

13 accepted papers

2025

Chain-of-Instructions: Compositional Instruction Tuning on Large Language Models

AAAI 2025technical

Fine-tuning large language models (LLMs) with a collection of large and diverse instructions has improved the model’s generalization to different tasks, even for unseen tasks. However, most existing instruction datasets include only single instructions, and they struggle to follow complex instructio…

2024

Generating Contextual Images for Long-Form Text

COLING 2024main

We investigate the problem of synthesizing relevant visual imagery from generic long-form text, leveraging Large Language Models (LLMs) and Text-to-Image Models (TIMs). Current Text-to-Image models require short prompts that describe the image content and style explicitly. Unlike image prompts, gene…

Cited by 0SourcePDFScholar
2020

Improving Device Directedness Classification of Utterances With Semantic Lexical Features

ICASSP 2020accepted

User interactions with personal assistants like Alexa, Google Home and Siri are typically initiated by a wake term or wake-word. Several personal assistants feature "follow-up" modes that allow users to make additional interactions without the need of a wakeword. For the system to only respond when…

Cited by 0SourceScholar
2019

Towards Better Confidence Estimation for Neural Models

ICASSP 2019accepted

In this work we focus on confidence modeling for neural network based text classification and sequence to sequence models in the context of Natural Language Understanding (NLU) tasks. For most applications, the confidence of a neural network model in it's output is computed as a function of the post…

Cited by 0SourceScholar
2017

End-to-end ASR-free keyword search from speech

ICASSP 2017accepted

End-to-end (E2E) systems have achieved competitive results compared to conventional hybrid hidden Markov model (HMM)-deep neural network based automatic speech recognition (ASR) systems. Such E2E systems are attractive due to the lack of dependence on alignments between input acoustic and output gra…

Cited by 0SourceScholar
2017

End-to-end speech recognition and keyword search on low-resource languages

ICASSP 2017accepted

In recent years, so-called, “end-to-end” speech recognition systems have emerged as viable alternatives to traditional ASR frameworks. Keyword search, localizing an orthographic query in a speech corpus, is typically performed by using automatic speech recognition (ASR) to generate an index. Previou…

Cited by 58SourceScholar
2017

Knowledge distillation across ensembles of multilingual models for low-resource languages

ICASSP 2017accepted

This paper investigates the effectiveness of knowledge distillation in the context of multilingual models. We show that with knowledge distillation, Long Short-Term Memory(LSTM) models can be used to train standard feed-forward Deep Neural Network (DNN) models for a variety of low-resource languages…

Cited by 0SourceScholar
2017

Network architectures for multilingual speech representation learning

ICASSP 2017accepted

Multilingual (ML) representations play a key role in building speech recognition systems for low resource languages. The IARPA sponsored BABEL program focuses on building speech recognition (ASR) and keyword search (KWS) systems in over 24 languages with limited training data. The most common mechan…

Cited by 0SourceScholar
2016

Semantic word embedding neural network language models for automatic speech recognition

ICASSP 2016accepted

Semantic word embeddings have become increasingly important in natural language processing tasks over the last few years. This popularity is due to their ability to easily capture rich semantic information through a distributed representation and the availability of fast and scalable algorithms for…

Cited by 0SourceScholar
2015

Bidirectional recurrent neural network language models for automatic speech recognition

ICASSP 2015accepted

Recurrent neural network language models have enjoyed great success in speech recognition, partially due to their ability to model longer-distance context than word n-gram models. In recurrent neural networks (RNNs), contextual information from past inputs is modeled with the help of recurrent conne…

Cited by 0SourceScholar
2015

Unnormalized exponential and neural network language models

ICASSP 2015accepted

Model M, an exponential class-based language model, and neural network language models (NNLM's) have outperformed word n-gram language models over a wide range of tasks. However, these gains come at the cost of vastly increased computation when calculating word probabilities. For both models, the bu…

Cited by 0SourceScholar