Advancing Streaming ASR with Chunk-wise Attention and Trans-chunk Selective State Spaces
This paper explores enhancing streaming speech recognition through the integration of chunk-wise attention and selective state space models (SSMs). The proposed framework replaces the quadratic complexity of attention-based context incorporation with a fully recurrent module based on selective SSMs.…