← Search

Kaushal Bhogale

3 accepted papers

2024

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages

ACL 2024findings

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these 7348 hours, 1639 hours have already been transcribed, with a…

2023

IndicSUPERB: A Speech Processing Universal Performance Benchmark for Indian Languages

AAAI 2023technical

A cornerstone in AI research has been the creation and adoption of standardized training and test datasets to earmark the progress of state-of-the-art models. A particularly successful example is the GLUE dataset for training and evaluating Natural Language Understanding (NLU) models for English. Th…

2022

Towards Efficient and Effective Self-Supervised Learning of Visual Representations

ECCV 2022poster

"Self-supervision has emerged as a propitious method for visual representation learning after the recent paradigm shift from handcrafted pretext tasks to instance-similarity based approaches. Most state-of-the-art methods enforce similarity between various augmentations of a given image, while some…