← Search

Martin Radfar

9 accepted papers

2023

End-to-End Spoken Language Understanding Using Joint CTC Loss and Self-Supervised, Pretrained Acoustic Encoders

ICASSP 2023accepted

It is challenging to extract semantic meanings directly from audio signals in spoken language understanding (SLU), due to the lack of textual information. Popular end-to-end (E2E) SLU models utilize sequence-to-sequence automatic speech recognition (ASR) models to extract textual embeddings as input…

Cited by 0SourceScholar
2023

Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers

ICML 2023poster

Streaming speech recognition architectures are employed for low-latency, real-time applications. Such architectures are often characterized by their causality. Causal architectures emit tokens at each frame, relying only on current and past signal, while non-causal models are exposed to a window of…

Cited by 9SourcePDFScholar
2022

A Neural Prosody Encoder for End-to-End Dialogue Act Classification

ICASSP 2022accepted

Dialogue act classification (DAC) is a critical task for spoken language understanding in dialogue systems. Prosodic features such as energy and pitch have been shown to be useful for DAC. Despite their importance, little research has explored neural approaches to integrate prosodic features into en…

Cited by 0SourceScholar
2022

Multi-Task RNN-T with Semantic Decoder for Streamable Spoken Language Understanding

ICASSP 2022accepted

End-to-end Spoken Language Understanding (E2E SLU) has attracted increasing interest due to its advantages of joint optimization and low latency when compared to traditionally cascaded pipelines. Existing E2E SLU models usually follow a two-stage configuration where an Automatic Speech Recognition (…

Cited by 0SourceScholar
2022

Tie Your Embeddings Down: Cross-Modal Latent Spaces for End-to-end Spoken Language Understanding

ICASSP 2022accepted

End-to-end (E2E) spoken language understanding (SLU) systems can infer the semantics of a spoken utterance directly from an audio signal. However, training an E2E system remains a challenge, largely due to the scarcity of paired audio-semantics data. In this paper, we consider an E2E system as a mul…

Cited by 0SourceScholar
2021

Encoding Syntactic Knowledge in Transformer Encoder for Intent Detection and Slot Filling

AAAI 2021technical

We propose a novel Transformer encoder-based architecture with syntactical knowledge encoded for intent detection and slot filling. Specifically, we encode syntactic knowledge into the Transformer encoder by jointly training it to predict syntactic parse ancestors and part-of-speech of each token vi…

Cited by 42SourcePDFScholar
2021

End-to-End Multi-Channel Transformer for Speech Recognition

ICASSP 2021accepted

Transformers are powerful neural architectures that allow integrating different modalities using attention mechanisms. In this paper, we leverage the neural transformer architectures for multi-channel speech recognition systems, where the spectral and spatial information collected from different mic…

Cited by 0SourceScholar