← Search

Gopala Krishna Anumanchipalli

9 accepted papers

2024

SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in Hubert

ICASSP 2024accepted

Data-driven unit discovery in self-supervised learning (SSL) of speech has embarked on a new era of spoken language processing. Yet, the discovered units often remain in phonetic space and speech units beyond phonemes are largely underexplored. Here, we demonstrate that a syllabic organization emerg…

Cited by 18SourceScholar
2024

Self-Supervised Audio-Visual Soundscape Stylization

ECCV 2024poster

"Speech sounds convey a great deal of information about the scenes, resulting in a variety of effects ranging from reverberation to additional ambient sounds. In this paper, we manipulate input speech to sound as though it was recorded within a different scene, given an audio-visual conditional exam…

Cited by 4SourcePDFScholar
2024

Self-Supervised Models of Speech Infer Universal Articulatory Kinematics

ICASSP 2024accepted

Self-Supervised Learning (SSL) based models of speech have shown remarkable performance on a range of downstream tasks. These state-of-the-art models have remained blackboxes, but many recent studies have begun “probing” models like HuBERT, to correlate their internal representations to different as…

Cited by 0SourceScholar
2024

Towards an Interpretable Representation of Speaker Identity via Perceptual Voice Qualities

ICASSP 2024accepted

Unlike other data modalities such as text and vision, speech does not lend itself to easy interpretation. While lay people can understand how to describe an image or sentence via perception, non-expert descriptions of speech often end at high-level demographic information, such as gender or age. In…

Cited by 0SourceScholar
2023

A Fast and Accurate Pitch Estimation Algorithm Based on the Pseudo Wigner-Ville Distribution

ICASSP 2023accepted

Estimation of fundamental frequency (F0) in voiced segments of speech signals, also known as pitch tracking, plays a crucial role in pitch synchronous speech analysis, speech synthesis, and speech manipulation. In this paper, we capitalize on the high time and frequency resolution of the pseudo Wign…

Cited by 0SourceScholar
2023

Articulation GAN: Unsupervised Modeling of Articulatory Learning

ICASSP 2023accepted

Generative deep neural networks are widely used for speech synthesis, but most existing models directly generate waveforms or spectral outputs. Humans, however, produce speech by controlling articulators, which results in the production of speech sounds through physical properties of sound propagati…

Cited by 0SourceScholar
2023

Articulatory Representation Learning via Joint Factor Analysis and Neural Matrix Factorization

ICASSP 2023accepted

Articulatory representation learning is the fundamental research in modeling neural speech production system. Our previous work has established a deep paradigm to decompose the articulatory kinematics data into gestures, which explicitly model the phonological and linguistic structure encoded with h…

Cited by 0SourceScholar
2023

Evidence of Vocal Tract Articulation in Self-Supervised Learning of Speech

ICASSP 2023accepted

Recent self-supervised learning (SSL) models have proven to learn rich representations of speech, which can readily be utilized by diverse downstream tasks. To understand such utilities, various analyses have been done for speech SSL models to reveal which and how information is encoded in the learn…

Cited by 0SourceScholar
2023

Speaker-Independent Acoustic-to-Articulatory Speech Inversion

ICASSP 2023accepted

To build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is a promising inversion target, since this space captures the mechanics of speech…

Cited by 39SourceScholar