← Search

Vishal Sunder

7 accepted papers

2026

IN-SYNC: ADAPTATION OF SPEECH AWARE LARGE LANGUAGE MODELS FOR ASR WITH WORD LEVEL TIMESTAMP PREDICTIONS

ICASSP 2026oral

Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is critical for applications such as captioning, media search, and…

Cited by 0SourcePDFScholar
2025

A Non-autoregressive Model for Joint STT and TTS

ICASSP 2025accepted

In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimodal framework capable of handling the speech and text modalities as input either individually or together. The proposed mo…

Cited by 0SourceScholar
2024

End-To-End Real Time Tracking of Children's Reading with Pointer Network

ICASSP 2024accepted

In this work, we explore how a real time reading tracker can be built efficiently for children’s voices. While previously proposed reading trackers focused on ASR-based cascaded approaches, we propose a fully end-to-end model making it less prone to lags in voice tracking. We employ a pointer networ…

Cited by 0SourceScholar
2023

End-to-End Word-Level Disfluency Detection and Classification in Children's Reading Assessment

ICASSP 2023accepted

Disfluency detection and classification on children’s speech has a great potential for teaching reading skills. Word-level assessment of children’s speech can help teachers to effectively gauge their students’ progress. Hence, we propose a novel attention-based model to perform word-level disfluency…

Cited by 0SourceScholar
2023

Fine-Grained Textual Knowledge Transfer to Improve RNN Transducers for Speech Recognition and Understanding

ICASSP 2023accepted

RNN Tranducer (RNN-T) technology is very popular for building deployable models for end-to-end (E2E) automatic speech recognition (ASR) and spoken language understanding (SLU). Since these are E2E models operating on speech directly, there remains a potential to improve their performance using purel…

Cited by 0SourceScholar
2022

Towards End-to-End Integration of Dialog History for Improved Spoken Language Understanding

ICASSP 2022accepted

Dialog history plays an important role in spoken language understanding (SLU) performance in a dialog system. For end-to-end (E2E) SLU, previous work has used dialog history in text form, which makes the model dependent on a cascaded automatic speech recognizer (ASR). This rescinds the benefits of a…

Cited by 0SourceScholar
2021

Handling Class Imbalance in Low-Resource Dialogue Systems by Combining Few-Shot Classification and Interpolation

ICASSP 2021accepted

Utterance classification performance in low-resource dialogue systems is constrained by an inevitably high degree of data imbalance in class labels. We present a new end-to-end pairwise learning framework that is designed specifically to tackle this phenomenon by inducing a few-shot classification c…

Cited by 0SourceScholar