← Search

Yash Jain

7 accepted papers

2026

Aurelius: Relation Aware Text-to-Audio Generation At Scale

ICLR 2026poster

We present Aurelius, a new framework that enables relation aware text-to-audio (TTA) generation research at scale. Given the lack of essential audio event and relation corpora, \emph{Aurelius} contributes a large-scale audio event corpus \emph{AudioEventSet} and another large-scale relation corpus \…

Cited by 0SourcecodeScholar
2025

Local Prompt Optimization

NAACL 2025short

In recent years, the use of prompts to guide the output of Large Language Models have increased dramatically. However, even the best of experts struggle to choose the correct words to stitch up a prompt for the desired task. To solve this, LLM driven prompt optimization emerged as an important probl…

Cited by 0SourcePDFScholar
2025

RiTTA: Modeling Event Relations in Text-to-Audio Generation

EMNLP 2025

Existing text-to-audio (TTA) generation methods have neither systematically explored audio event relation modeling, nor proposed any new framework to enhance this capability. In this work, we systematically study audio event relation modeling in TTA generation models. We first establish a benchmark

2024

Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition

COLING 2024main

Recent advances in machine learning have demonstrated that multi-modal pre-training can improve automatic speech recognition (ASR) performance compared to randomly initialized models, even when models are fine-tuned on uni-modal tasks. Existing multi-modal pre-training methods for the ASR task have…

Cited by 2SourcePDFScholar
2024

PEEKABOO: Interactive Video Generation via Masked-Diffusion

CVPR 2024poster

Modern video generation models like Sora have achieved remarkable success in producing high-quality videos. However a significant limitation is their inability to offer interactive control to users a feature that promises to open up unprecedented applications and creativity. In this work we introduc…

2023

DAMEX: Dataset-aware Mixture-of-Experts for visual understanding of mixture-of-datasets

NeurIPS 2023poster

Construction of a universal detector poses a crucial question: How can we most effectively train a model on a large mixture of datasets? The answer lies in learning dataset-specific features and ensembling their knowledge but do all this in a single model. Previous methods achieve this by h…

2021

Speech Recognition Using RFID Tattoos (Extended Abstract)

IJCAI 2021poster

This paper presents a radio-frequency (RF) based assistive technology for voice impairments (i.e., dysphonia), which occurs in an estimated 1% of the global population. We specifically focus on acquired voice disorders where users continue to be able to make facial and lip gestures associated with s…

Cited by 0SourcePDFScholar