← Search

Piotr Zelasko

10 accepted papers

2025

Chain-of-Thought Prompting for Speech Translation

ICASSP 2025accepted

Large language models (LLMs) have demonstrated remarkable advancements in language understanding and generation. Building on the success of text-based LLMs, recent research has adapted these models to use speech embeddings for prompting, resulting in Speech-LLM models that exhibit strong performance…

Cited by 0SourceScholar
2025

EMMeTT: Efficient Multimodal Machine Translation Training

ICASSP 2025accepted

A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses on neural machine translation (NMT) and proposes a joint multimodal training regime of Speech-LLM to include automatic s…

Cited by 0SourceScholar
2025

VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning

NAACL 2025long

Recent studies have augmented large language models (LLMs) with speech capabilities, leading to the development of speech language models (SpeechLMs). Earlier SpeechLMs focused on single-turn speech-based question answering (QA), where user input comprised a speech context and a text question. More…

2023

Delay-Penalized Transducer for Low-Latency Streaming ASR

ICASSP 2023accepted

In streaming automatic speech recognition (ASR), it is desirable to reduce latency as much as possible while having minimum impact on recognition accuracy. Although a few existing methods are able to achieve this goal, they are difficult to implement due to their dependency on external alignments. I…

Cited by 0SourceScholar
2023

Fast and Parallel Decoding for Transducer

ICASSP 2023accepted

The transducer architecture is becoming increasingly popular in the field of speech recognition, because it is naturally streaming as well as high in accuracy. One of the drawbacks of transducer is that it is difficult to decode in a fast and parallel way due to an unconstrained number of symbols th…

Cited by 0SourceScholar
2023

Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation

ICASSP 2023accepted

Knowledge distillation (KD) is a common approach to improve model performance in automatic speech recognition (ASR), where a student model is trained to imitate the output behaviour of a teacher model. However, traditional KD methods suffer from teacher label storage issue, especially when the train…

Cited by 0SourceScholar
2021

CopyPaste: An Augmentation Method for Speech Emotion Recognition

ICASSP 2021accepted

Data augmentation is a widely used strategy for training robust machine learning models. It partially alleviates the problem of limited data for tasks like speech emotion recognition (SER), where collecting data is expensive and challenging. This study proposes CopyPaste, a perceptually motivated no…

Cited by 0SourceScholar
2021

Focus on the Present: A Regularization Method for the ASR Source-Target Attention Layer

ICASSP 2021accepted

This paper introduces a novel method to diagnose the source-target attention in state-of-the-art end-to-end speech recognition models with joint connectionist temporal classification (CTC) and attention training. Our method is based on the fact that both, CTC and source-target attention, are acting…

Cited by 0SourceScholar
2021

How Phonotactics Affect Multilingual and Zero-Shot ASR Performance

ICASSP 2021accepted

The idea of combining multiple languages’ recordings to train a single automatic speech recognition (ASR) model brings the promise of the emergence of universal speech representation. Recently, a Transformer encoder-decoder model has been shown to leverage multilingual data well in IPA transcription…

Cited by 0SourceScholar
2021

Improving Reconstruction Loss Based Speaker Embedding in Unsupervised and Semi-Supervised Scenarios

ICASSP 2021accepted

Text-to-speech (TTS) models trained to minimize the spectrogram reconstruction loss can learn speaker embeddings without explicit speaker identity supervision, unlike x-vector speaker identification (SID) systems. Leveraging this way of speaker embedding learning can be useful in unsupervised or sem…

Cited by 0SourceScholar