← Search

Jan Chorowski

6 accepted papers

2023

Efficient Transformers with Dynamic Token Pooling

ACL 2023long

Transformers achieve unrivalled performance in modelling language, but remain inefficient in terms of memory and time complexity. A possible remedy is to reduce the sequence length in the intermediate layers by pooling fixed-length segments of tokens. Nevertheless, natural units of meaning, such as…

2022

Contrastive Prediction Strategies for Unsupervised Segmentation and Categorization of Phonemes and Words

ICASSP 2022accepted

We identify a performance trade-off between the tasks of phoneme categorization and phoneme and word segmentation in several self-supervised learning algorithms based on Contrastive Predictive Coding (CPC). Our experiments suggest that context building networks, albeit necessary for high performance…

Cited by 0SourceScholar
2018

On Using Backpropagation for Speech Texture Generation and Voice Conversion

ICASSP 2018accepted

Inspired by recent work on neural network image generation which rely on backpropagation towards the network inputs, we present a proof-of-concept system for speech texture synthesis and voice conversion based on two mechanisms: approximate inversion of the representation learned by a speech recogni…

Cited by 0SourceScholar
2018

State-of-the-Art Speech Recognition with Sequence-to-Sequence Models

ICASSP 2018accepted

Attention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a single neural network. In previous work, we have shown that such architectures ar…

Cited by 0SourceScholar
2017

Input Switched Affine Networks: An RNN Architecture Designed for Interpretability

ICML 2017poster

There exist many problem domains where the interpretability of neural network models is essential for deployment. Here we introduce a recurrent architecture composed of input-switched affine transformations – in other words an RNN without any explicit nonlinearities, but with input-dependent recurre…

Cited by 41SourcePDFScholar
2016

End-to-end attention-based large vocabulary speech recognition

ICASSP 2016accepted

Many state-of-the-art Large Vocabulary Continuous Speech Recognition (LVCSR) Systems are hybrids of neural networks and Hidden Markov Models (HMMs). Recently, more direct end-to-end methods have been investigated, in which neural architectures were trained to model sequences of characters [1,2]. To…

Cited by 0SourceScholar