← Search

Rami Botros

4 accepted papers

2023

Lego-Features: Exporting Modular Encoder Features for Streaming and Deliberation ASR

ICASSP 2023accepted

In end-to-end (E2E) speech recognition models, a representational tight-coupling inevitably emerges between the encoder and the decoder. We build upon recent work that has begun to explore building encoders with modular encoded representations, such that encoders and decoders from different models c…

Cited by 3SourceScholar
2022

Improving The Latency And Quality Of Cascaded Encoders

ICASSP 2022accepted

In this paper, we explore reducing computational latency of the 2-pass cascaded encoder model [1]. Specifically, we experiment with reducing the size of the causal 1st-pass and adding capacity to the non-causal 2nd-pass, such that the overall latency can be reduced without loss of quality. In additi…

Cited by 0SourceScholar
2019

A Neural Network Based Ranking Framework to Improve ASR with NLU Related Knowledge Deployed

ICASSP 2019accepted

This work proposes a new neural network framework to simultaneously rank multiple hypotheses generated by one or more automatic speech recognition (ASR) engines for a speech utterance. Features fed in the framework not only include those calculated from the ASR information, but also involve natural…

Cited by 0SourceScholar
2017

A deep learning approach to traffic lights: Detection, tracking, and classification

ICRA 2017poster

Reliable traffic light detection and classification is crucial for automated driving in urban environments. Currently, there are no systems that can reliably perceive traffic lights in real-time, without map-based information, and in sufficient distances needed for smooth urban driving. We propose a…

Cited by 350SourceScholar