← Search

Dmitriy Serdyuk

6 accepted papers

2024

Conformer is All You Need for Visual Speech Recognition

ICASSP 2024accepted

Visual speech recognition models extract visual features in a hierarchical manner. At the lower level, there is a visual front-end with a limited temporal receptive field that processes the raw pixels depicting the lips or faces. At the higher level, there is an encoder that attends to the embedding…

Cited by 0SourceScholar
2018

Deep Complex Networks

ICLR 2018poster

At present, the vast majority of building blocks, techniques, and architectures for deep learning are based on real-valued operations and representations. However, recent work on recurrent neural networks and older fundamental theoretical analysis suggests that complex numbers could have a richer re…

2018

Towards End-to-end Spoken Language Understanding

ICASSP 2018accepted

Spoken language understanding system is traditionally designed as a pipeline of a number of components. First, the audio signal is processed by an automatic speech recognizer for transcription or n-best hypotheses. With the recognition results, a natural language understanding system classifies the…

Cited by 0SourceScholar
2018

Twin Networks: Matching the Future for Sequence Generation

ICLR 2018poster

We propose a simple technique for encouraging generative RNNs to plan ahead. We train a ``backward'' recurrent network to generate a given sequence in reverse order, and we encourage states of the forward model to predict cotemporal states of the backward model. The backward network is used only dur…

2016

End-to-end attention-based large vocabulary speech recognition

ICASSP 2016accepted

Many state-of-the-art Large Vocabulary Continuous Speech Recognition (LVCSR) Systems are hybrids of neural networks and Hidden Markov Models (HMMs). Recently, more direct end-to-end methods have been investigated, in which neural architectures were trained to model sequences of characters [1,2]. To…

Cited by 0SourceScholar
2015

Attention-Based Models for Speech Recognition

NeurIPS 2015spotlight

Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks including machine translation, handwriting synthesis and image caption generation. We extend the attention-mechanism with features needed for speech re…

Cited by 3496SourcePDFScholar