Lego-Features: Exporting Modular Encoder Features for Streaming and Deliberation ASR
In end-to-end (E2E) speech recognition models, a representational tight-coupling inevitably emerges between the encoder and the decoder. We build upon recent work that has begun to explore building encoders with modular encoded representations, such that encoders and decoders from different models c…