← Search

Alan W Black

27 accepted papers

2026

Scaling Spoken Language Models with Syllabic Speech Tokenization

ICASSP 2026oral

Spoken language models (SLMs) typically discretize speech into high-frame-rate tokens extracted from SSL speech models. As the most successful LMs are based on the Transformer architecture, processing these long token streams with self-attention is expensive, as attention scales quadratically with s…

Cited by 0SourcePDFScholar
2024

SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in Hubert

ICASSP 2024accepted

Data-driven unit discovery in self-supervised learning (SSL) of speech has embarked on a new era of spoken language processing. Yet, the discovered units often remain in phonetic space and speech units beyond phonemes are largely underexplored. Here, we demonstrate that a syllabic organization emerg…

Cited by 0SourceScholar
2024

Self-Supervised Models of Speech Infer Universal Articulatory Kinematics

ICASSP 2024accepted

Self-Supervised Learning (SSL) based models of speech have shown remarkable performance on a range of downstream tasks. These state-of-the-art models have remained blackboxes, but many recent studies have begun “probing” models like HuBERT, to correlate their internal representations to different as…

Cited by 0SourceScholar
2023

A Fast and Accurate Pitch Estimation Algorithm Based on the Pseudo Wigner-Ville Distribution

ICASSP 2023accepted

Estimation of fundamental frequency (F0) in voiced segments of speech signals, also known as pitch tracking, plays a crucial role in pitch synchronous speech analysis, speech synthesis, and speech manipulation. In this paper, we capitalize on the high time and frequency resolution of the pseudo Wign…

Cited by 0SourceScholar
2023

Articulatory Representation Learning via Joint Factor Analysis and Neural Matrix Factorization

ICASSP 2023accepted

Articulatory representation learning is the fundamental research in modeling neural speech production system. Our previous work has established a deep paradigm to decompose the articulatory kinematics data into gestures, which explicitly model the phonological and linguistic structure encoded with h…

Cited by 0SourceScholar
2023

Speaker-Independent Acoustic-to-Articulatory Speech Inversion

ICASSP 2023accepted

To build speech processing methods that can handle speech as naturally as humans, researchers have explored multiple ways of building an invertible mapping from speech to an interpretable space. The articulatory space is a promising inversion target, since this space captures the mechanics of speech…

Cited by 0SourceScholar
2022

ESPnet-SLU: Advancing Spoken Language Understanding Through ESPnet

ICASSP 2022accepted

As Automatic Speech Processing (ASR) systems are getting better, there is an increasing interest of using the ASR output to do downstream Natural Language Processing (NLP) tasks. However, there are few open source toolkits that can be used to generate reproducible results on different Spoken Languag…

Cited by 0SourceScholar
2022

End-to-End Speech Summarization Using Restricted Self-Attention

ICASSP 2022accepted

Speech summarization is typically performed by using a cascade of speech recognition and text summarization models. End-to-end modeling of speech summarization models is challenging due to memory and compute constraints arising from long input audio sequences. Recent work in document summarization h…

Cited by 0SourceScholar
2022

On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization

EMNLP 2022finding

Combining the visual modality with pretrained language models has been surprisingly effective for simple descriptive tasks such as image captioning. More general text generation however remains elusive. We take a step back and ask: How do these models work for more complex generative tasks, i.e. con…

Cited by 3SourcePDFScholar
2022

Token-level Sequence Labeling for Spoken Language Understanding using Compositional End-to-End Models

EMNLP 2022finding

End-to-end spoken language understanding (SLU) systems are gaining popularity over cascaded approaches due to their simplicity and ability to avoid error propagation. However, these systems model sequence labeling as a sequence prediction task causing a divergence from its well-established token-lev…

2021

Acoustics Based Intent Recognition Using Discovered Phonetic Units for Low Resource Languages

ICASSP 2021accepted

With recent advancements in language technologies, humans are now speaking to devices. Increasing the reach of spoken language technologies requires building systems in local languages. A major bottleneck here are the underlying data-intensive parts that make up such systems, including automatic spe…

Cited by 0SourceScholar
2021

Breaking Down Walls of Text: How Can NLP Benefit Consumer Privacy?

ACL 2021long

Privacy plays a crucial role in preserving democratic ideals and personal autonomy. The dominant legal approach to privacy in many jurisdictions is the “Notice and Choice” paradigm, where privacy policies are the primary instrument used to convey information to users. However, privacy policies are l…

Cited by 34SourcePDFScholar
2021

Focused Attention Improves Document-Grounded Generation

NAACL 2021long

Document grounded generation is the task of using the information provided in a document to improve text generation. This work focuses on two different document grounded generation tasks: Wikipedia Update Generation task and Dialogue response generation. Our work introduces two novel adaptations of…

2021

Multilingual Phonetic Dataset for Low Resource Speech Recognition

ICASSP 2021accepted

Phone Recognition is one of the most important tasks in the field of multilingual speech recognition, especially for low-resource languages whose orthographies are not available. However, most speech recognition datasets so far only focus on high-resource languages, there are very few datasets avail…

Cited by 0SourceScholar
2021

Phone Distribution Estimation for Low Resource Languages

ICASSP 2021accepted

Phones are critical components in various computational linguistic fields, for example, phone distributions could be helpful in speech recognition and speech synthesis. Traditional approaches to estimate phone distributions typically involve G2P systems which are either manually designed by linguist…

Cited by 0SourceScholar
2021

Switch Point biased Self-Training: Re-purposing Pretrained Models for Code-Switching

EMNLP 2021finding

Code-switching (CS), a ubiquitous phenomenon due to the ease of communication it offers in multilingual communities still remains an understudied problem in language processing. The primary reasons behind this are: (1) minimal efforts in leveraging large pretrained multilingual models, and (2) the l…

Cited by 5SourcePDFScholar
2020

Augmenting Non-Collaborative Dialog Systems with Explicit Semantic and Strategic Dialog History

ICLR 2020poster

We study non-collaborative dialogs, where two agents have a conflict of interest but must strategically communicate to reach an agreement (e.g., negotiation). This setting poses new challenges for modeling dialog history because the dialog's outcome relies not only on the semantic intent, but also o…

Cited by 37SourceScholar
2020

Universal Phone Recognition with a Multilingual Allophone System

ICASSP 2020accepted

Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can support lexical contrasts in a particular language) and thei…

Cited by 0SourceScholar
2019

CMU Wilderness Multilingual Speech Dataset

ICASSP 2019accepted

This paper describes the CMU Wilderness Multilingual Speech Dataset. A dataset of over 700 different languages providing audio, aligned text and word pronunciations. On average each language provides around 20 hours of sentence-lengthed transcriptions. We describe our multi-pass alignment techniques…

Cited by 0SourceScholar
2019

Learning Disentangled Representation in Latent Stochastic Models: A Case Study with Image Captioning

ICASSP 2019accepted

Multimodal tasks require learning joint representation across modalities. In this paper, we present an approach to employ latent stochastic models for a multimodal task image captioning. Encoder Decoder models with stochastic latent variables are often faced with optimization issues such as latent c…

Cited by 0SourceScholar
2019

Phoneme Level Language Models for Sequence Based Low Resource ASR

ICASSP 2019accepted

Building multilingual and crosslingual models help bring different languages together in a language universal space. It allows models to share parameters and transfer knowledge across languages, enabling faster and better adaptation to a new language. These approaches are particularly useful for low…

Cited by 0SourceScholar
2018

Linguistic Unit Discovery from Multi-Modal Inputs in Unwritten Languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop

ICASSP 2018accepted

We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and words) in a language without orthography. We study the replacement of orthographic transcriptions by images and/or translate…

Cited by 0SourceScholar
2018

Sequence-Based Multi-Lingual Low Resource Speech Recognition

ICASSP 2018accepted

Techniques for multi-lingual and cross-lingual speech recognition can help in low resource scenarios, to bootstrap systems and enable analysis of new languages and domains. End-to-end approaches, in particular sequence-based techniques, are attractive because of their simplicity and elegance. While…

Cited by 0SourceScholar
2015

Modulation spectrum-constrained trajectory training algorithm for GMM-based Voice Conversion

ICASSP 2015accepted

This paper presents a novel training algorithm for Gaussian Mixture Model (GMM)-based Voice Conversion (VC). One of the advantages of GMM-based VC is computationally efficient conversion processing enabling to achieve real-time VC applications. On the other hand, the quality of the converted speech…

Cited by 0SourceScholar
2015

Parameter generation algorithm considering Modulation Spectrum for HMM-based speech synthesis

ICASSP 2015accepted

This paper proposes a novel parameter generation algorithm for high-quality speech generation in Hidden Markov Model (HMM)-based speech synthesis. One of the biggest issues causing significant quality degradation is the over-smoothing effect often observed in generated parameter trajectories. Global…

Cited by 0SourceScholar