← Search

K. Sri Rama Murty

7 accepted papers

2025

Stream-TTS: A Low-Latency Text-to-Speech using Kolmogorov-Arnold Networks for Streaming Speech Applications

ICASSP 2025accepted

The rise of conversational AI and multimodal streaming applications has led to a significant demand for low-latency Text-to-Speech (TTS) systems. This work presents a multilingual low-latency model that leverages the functional decomposition principles of Kolmogorov-Arnold Networks (KANs) in modelin…

Cited by 0SourceScholar
2023

Lightweight Prosody-TTS for Multi-Lingual Multi-Speaker Scenario

ICASSP 2023accepted

This work presents a lightweight end-to-end text-to-speech (TTS) synthesis for the multi-lingual multi-speaker (ML-MS) scenario. The proposed system uses nonautoregressive modular architecture with interconnected subnets for text-encoder, duration estimator, f<inf xmlns:mml="http://www.w3.org/1998/M…

Cited by 0SourceScholar
2022

Multi-Feature Integration for Speaker Embedding Extraction

ICASSP 2022accepted

The performance of the automatic speaker recognition system is becoming more and more accurate, with the advancement in deep learning methods. However, current speaker recognition system performances are subjective to the training conditions, thereby decreasing the performance drastically even on sl…

Cited by 0SourceScholar
2019

Importance of Analytic Phase of the Speech Signal for Detecting Replay Attacks in Automatic Speaker Verification Systems

ICASSP 2019accepted

In this paper, the importance of analytic phase of the speech signal in automatic speaker verification systems is demonstrated in the context of replay spoof attacks. In order to accurately detect the replay spoof attacks, effective feature representations of speech signals are required to capture t…

Cited by 0SourceScholar
2019

Zero Resource Speaking Rate Estimation from Change Point Detection of Syllable-like Units

ICASSP 2019accepted

Speaking rate is an important attribute of the speech signal which plays a crucial role in the performance of automatic speech processing systems. In this paper, we propose to estimate the speaking rate by segmenting the speech into syllable-like units using end point detection algorithms which do n…

Cited by 0SourceScholar
2018

Phoneme Based Embedded Segmental K-Means for Unsupervised Term Discovery

ICASSP 2018accepted

Identifying and grouping the frequently occurring word-like patterns from raw acoustic waveforms is an important task in the zero resource speech processing. Embedded segmental K-means (ES-KMeans) discovers both the word boundaries and the word types from raw data. Starting from an initial set of su…

Cited by 0SourceScholar
2017

Action-vectors: Unsupervised movement modeling for action recognition

ICASSP 2017accepted

Representation and modelling of movements play a significant role in recognising actions in unconstrained videos. However, explicit segmentation and labelling of movements are non-trivial because of the variability associated with actors, camera viewpoints, duration etc. Therefore, we propose to tra…

Cited by 0SourceScholar