← Search

Ashish Shrivastava

10 accepted papers

2023

SHIFT3D: Synthesizing Hard Inputs For Tricking 3D Detectors

ICCV 2023poster

We present SHIFT3D, a differentiable pipeline for generating 3D shapes that are structurally plausible yet challenging to 3D object detectors. In safety-critical applications like autonomous driving, discovering such novel challenging objects can offer insight into unknown vulnerabilities of 3D dete…

Cited by 1PDFScholar
2022

SYNT++: Utilizing Imperfect Synthetic Data to Improve Speech Recognition

ICASSP 2022accepted

With recent advances in speech synthesis, synthetic data is becoming a viable alternative to real data for training speech recognition models. However, machine learning with synthetic data is not trivial due to the gap between the synthetic and the real data distributions. Synthetic datasets may con…

Cited by 0SourceScholar
2022

Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models

ICML 2022spotlight

Controllable generative sequence models with the capability to extract and replicate the style of specific examples enable many applications, including narrating audiobooks in different voices, auto-completing and auto-correcting written handwriting, and generating missing training samples for downs…

Cited by 26SourcePDFScholar
2021

Optimize What Matters: Training DNN-Hmm Keyword Spotting Model Using End Metric

ICASSP 2021accepted

Deep Neural Network–Hidden Markov Model (DNN-HMM) based methods have been successfully used for many always-on keyword spotting algorithms that detect a wake word to trigger a device. The DNN predicts the state probabilities of a given speech frame, while HMM decoder combines the DNN predictions of…

Cited by 0SourceScholar
2021

SLAP: a Split Latency Adaptive VLIW Pipeline Architecture Which Enables on-The-Fly Variable SIMD Vector-Length

ICASSP 2021accepted

Over the last decade the relative latency of access to shared memory by multicore increased as wire resistance dominated latency and low wire density layout pushed multi-port memories farther away from their ports. Various techniques were deployed to improve average memory access latencies, such as…

Cited by 0SourceScholar
2021

SapAugment: Learning A Sample Adaptive Policy for Data Augmentation

ICASSP 2021accepted

Data augmentation methods usually apply the same augmentation (or a mix of them) to all the training samples. For example, to perturb data with noise, the noise is sampled from a Normal distribution with a fixed standard deviation, for all samples. We hypothesize that a hard sample with high trainin…

Cited by 0SourceScholar
2021

Saying No is An Art: Contextualized Fallback Responses for Unanswerable Dialogue Queries

ACL 2021short

Despite end-to-end neural systems making significant progress in the last decade for task-oriented as well as chit-chat based dialogue systems, most dialogue systems rely on hybrid approaches which use a combination of rule-based, retrieval and generative approaches for generating a set of ranked re…

2020

Unsupervised Style and Content Separation by Minimizing Mutual Information for Speech Synthesis

ICASSP 2020accepted

We present a method to generate speech from input text and a style vector that is extracted from a reference speech signal in an unsupervised manner, i.e., no style annotation, such as speaker information, is required. Existing unsupervised methods, during training, generate speech by computing styl…

Cited by 0SourceScholar
2017

Learning From Simulated and Unsupervised Images Through Adversarial Training

CVPR 2017oral

With recent progress in graphics, it has become more tractable to train models on synthetic images, potentially avoiding the need for expensive annotations. However, learning from synthetic images may not achieve the desired performance due to a gap between synthetic and real image distributions. To…

Cited by 2368PDFScholar
2015

Class Consistent Multi-Modal Fusion With Binary Features

CVPR 2015poster

Many existing recognition algorithms combine different modalities based on training accuracy but do not consider the possibility of noise at test time. We describe an algorithm that perturbs test features so that all modalities predict the same class. We enforce this perturbation to be as small as p…

Cited by 15SourcePDFScholar