← Search

Pratyush Kumar

19 accepted papers

2025

Towards Bringing Parity in Pretraining Datasets for Low-resource Indian Languages

ICASSP 2025accepted

Lack of large-scale pretraining data for low resource languages from the Indian sub-continent, leads to their underrepresentation in existing massively multilingual models. In this work, we address this gap by proposing a framework to create large raw audio datasets for such under-represented langua…

Cited by 0SourceScholar
2024

IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages

ACL 2024long

Despite the considerable advancements in English LLMs, the progress in building comparable models for other languages has been hindered due to the scarcity of tailored resources. Our work aims to bridge this divide by introducing an expansive suite of resources specifically designed for the developm…

2024

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages

ACL 2024findings

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these 7348 hours, 1639 hours have already been transcribed, with a…

2023

Effectiveness of Mining Audio and Text Pairs from Public Data for Improving ASR Systems for Low-Resource Languages

ICASSP 2023accepted

Collecting labelled datasets for speech recognition systems for low-resource languages on a diverse set of domains and speakers is expensive. In this work, we demonstrate an inexpensive and effective alternative by "mining" text and audio pairs for Indian languages from public sources, specifically…

Cited by 0SourceScholar
2023

IndicMT Eval: A Dataset to Meta-Evaluate Machine Translation Metrics for Indian Languages

ACL 2023long

The rapid growth of machine translation (MT) systems necessitates meta-evaluations of evaluation metrics to enable selection of those that best reflect MT quality. Unfortunately, most meta-evaluation studies focus on European languages, the observations for which may not always apply to other langua…

2023

IndicSUPERB: A Speech Processing Universal Performance Benchmark for Indian Languages

AAAI 2023technical

A cornerstone in AI research has been the creation and adoption of standardized training and test datasets to earmark the progress of state-of-the-art models. A particularly successful example is the GLUE dataset for training and evaluating Natural Language Understanding (NLU) models for English. Th…

2023

Naamapadam: A Large-Scale Named Entity Annotated Data for Indic Languages

ACL 2023long

We present, Naamapadam, the largest publicly available Named Entity Recognition (NER) dataset for the 11 major Indian languages from two language families. The dataset contains more than 400k sentences annotated with a total of at least 100k entities from three standard entity categories (Person, Lo…

2023

Towards Building Text-to-Speech Systems for the Next Billion Users

ICASSP 2023accepted

Deep learning based text-to-speech (TTS) systems have been evolving rapidly with advances in model architectures, training methodologies, and generalization across speakers and languages. However, these advances have not been thoroughly investigated for Indian language speech synthesis. Such investi…

Cited by 0SourceScholar
2023

Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages

ACL 2023long

Building Natural Language Understanding (NLU) capabilities for Indic languages, which have a collective speaker base of more than one billion speakers is absolutely crucial. In this work, we aim to improve the NLU capabilities of Indic languages by making contributions along 3 important axes (i) mon…

2022

Addressing Resource Scarcity across Sign Languages with Multilingual Pretraining and Unified-Vocabulary Datasets

NeurIPS 2022accept

There are over 300 sign languages in the world, many of which have very limited or no labelled sign-to-text datasets. To address low-resource data scenarios, self-supervised pretraining and multilingual finetuning have been shown to be effective in natural language and speech processing. In this wor…

Cited by 10SourcePDFScholar
2022

IndicBART: A Pre-trained Model for Indic Natural Language Generation

ACL 2022findings

In this paper, we study pre-trained sequence-to-sequence models for a group of related languages, with a focus on Indic languages. We present IndicBART, a multilingual, sequence-to-sequence pre-trained model focusing on 11 Indic languages and English. IndicBART utilizes the orthographic similarity b…

2022

IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages

EMNLP 2022main

Natural Language Generation (NLG) for non-English languages is hampered by the scarcity of datasets in these languages. We present the IndicNLG Benchmark, a collection of datasets for benchmarking NLG for 11 Indic languages. We focus on five diverse tasks, namely, biography generation using Wikipedi…

2022

Input-specific Attention Subnetworks for Adversarial Detection

ACL 2022findings

Self-attention heads are characteristic of Transformer models and have been well studied for interpretability and pruning. In this work, we demonstrate an altogether different utility of attention heads, namely for adversarial detection. Specifically, we propose a method to construct input-specific…

2022

OpenHands: Making Sign Language Recognition Accessible with Pose-based Pretrained Models across Languages

ACL 2022long

AI technologies for Natural Languages have made tremendous progress recently. However, commensurate progress has not been made on Sign Languages, in particular, in recognizing signs as individual words or as complete sentences. We introduce OpenHands, a library where we take four key ideas from the…

Cited by 59SourcePDFScholar
2022

Towards Building ASR Systems for the Next Billion Users

AAAI 2022technical

Recent methods in speech and language technology pretrain very large models which are fine-tuned for specific tasks. However, the benefits of such large models are often limited to a few resource rich languages of the world. In this work, we make multiple contributions towards building ASR systems f…

2021

A Systematic Evaluation of Object Detection Networks for Scientific Plots

AAAI 2021technical

Are existing object detection methods adequate for detecting text and visual elements in scientific plots which are arguably different than the objects found in natural images? To answer this question, we train and compare the accuracy of Fast/Faster R-CNN, SSD, YOLO and RetinaNet on the PlotQA data…

Cited by 9SourcePDFScholar
2021

The Heads Hypothesis: A Unifying Statistical Approach Towards Understanding Multi-Headed Attention in BERT

AAAI 2021technical

Multi-headed attention heads are a mainstay in transformer-based models. Different methods have been proposed to classify the role of each attention head based on the relations between tokens which have high pair-wise attention. These roles include syntactic (tokens with some syntactic relation), lo…

2020

Joint Transformer/RNN Architecture for Gesture Typing in Indic Languages

COLING 2020main

Gesture typing is a method of typing words on a touch-based keyboard by creating a continuous trace passing through the relevant keys. This work is aimed at developing a keyboard that supports gesture typing in Indic languages. We begin by noting that when dealing with Indic languages, one needs to…

2018

Opportunistic Sensing with MIC Arrays on Smart Speakers for Distal Interaction and Exercise Tracking

ICASSP 2018accepted

In 2017, smart speakers (such as Amazon Echo, Google Home, etc.) became a commercial success. Most smart speakers have a circular microphone array to provide hands-free, voice-only interaction from a distance. In this work, we exploit this mic array for opportunistically sensing gestures and trackin…

Cited by 0SourceScholar