← Search

Julian Salazar

7 accepted papers

2025

Long-Form Speech Generation with Spoken Language Models

ICML 2025oral

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, textless spoken language models struggle to generate plausible speech past tens of seconds, due to high temporal resolution of speech tok…

2025

Prompting with Phonemes: Enhancing LLMs’ Multilinguality for Non-Latin Script Languages

NAACL 2025long

Multilingual LLMs have achieved remarkable benchmark performance, but we find they continue to underperform on non-Latin script languages across contemporary LLM families. This discrepancy arises from the fact that LLMs are pretrained with orthographic scripts, which are dominated by Latin character…

Cited by 0SourcePDFScholar
2025

Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech

NAACL 2025long

Autoregressive (AR) Transformer-based sequence models are known to have difficulty generalizing to sequences longer than those seen during training. When applied to text-to-speech (TTS), these models tend to drop or repeat words or produce erratic output, especially for longer utterances. In this pa…

2024

Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

ICLR 2024poster

We present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire sy…

Cited by 40SourcePDFScholar
2021

Align-Refine: Non-Autoregressive Speech Recognition via Iterative Realignment

NAACL 2021long

Non-autoregressive encoder-decoder models greatly improve decoding speed over autoregressive models, at the expense of generation quality. To mitigate this, iterative decoding models repeatedly infill or refine the proposal of a non-autoregressive model. However, editing at the level of output seque…

2020

Deep Contextualized Acoustic Representations for Semi-Supervised Speech Recognition

ICASSP 2020accepted

We propose a novel approach to semi-supervised automatic speech recognition (ASR). We first exploit a large amount of unlabeled audio data via representation learning, where we reconstruct a temporal slice of filterbank features from past and future context frames. The resulting deep contextualized…

Cited by 0SourceScholar
2019

Self-attention Networks for Connectionist Temporal Classification in Speech Recognition

ICASSP 2019accepted

The success of self-attention in NLP has led to recent applications in end-to-end encoder-decoder architectures for speech recognition. Separately, connectionist temporal classification (CTC) has matured as an alignment-free, non-autoregressive approach to sequence transduction, either by itself or…

Cited by 0SourceScholar