← Search

Srikanth Vishnubhotla

6 accepted papers

2025

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation

ICASSP 2025accepted

Audio-Visual Speech-to-Speech Translation (AVS2S) typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony—ensuring that the movements of the lips match the spoken content—essential for maintaining realism in du…

Cited by 0SourceScholar
2025

Knowledge Distillation From Ensemble for Spoken Language Identification

ICASSP 2025accepted

Spoken language identification (LID) has seen substantial performance gains with the rise of large-scale models. However, these models are often computationally expensive and impractical for many real-world applications. In this work, we propose a novel knowledge distillation from ensemble framework…

Cited by 0SourceScholar
2025

MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization

EMNLP 2025

Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods are essential for reducing both cost and human effort. However, such methods require a strong meta-evaluation benchmark g

Cited by 0SourcePDFScholar
2024

SpeechGuard: Exploring the Adversarial Robustness of Multi-modal Large Language Models

ACL 2024findings

Integrated Speech and Large Language Models (SLMs) that can follow speech instructions and generate relevant text responses have gained popularity lately. However, the safety and robustness of these models remains largely unclear. In this work, we investigate the potential vulnerabilities of such in…

2021

Knowledge Transfer for Efficient on-Device False Trigger Mitigation

ICASSP 2021accepted

In this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as a "false trigger" and false trigger mitigation (FTM) is essential for designing a privacy-centric non-intrusive smart as…

Cited by 0SourceScholar
2020

Lattice-Based Improvements for Voice Triggering Using Graph Neural Networks

ICASSP 2020accepted

Voice-triggered smart assistants often rely on detection of a trigger-phrase before they start listening for the user request. Mitigation of false triggers is an important aspect of building a privacy-centric non-intrusive smart assistant. In this paper, we address the task of false trigger mitigati…

Cited by 0SourceScholar