← Search

Grant Strimel

4 accepted papers

2025

Context-aware Dynamic Pruning for Speech Foundation Models

ICLR 2025poster

Foundation models, such as large language models, have achieved remarkable success in natural language processing and are evolving into models capable of handling multiple modalities. Listening ability, in particular, is crucial for many applications, leading to research on building speech foundatio…

Cited by 0SourcePDFScholar
2025

SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning

ACL 2025long

We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50M is built from publicly available speech corpora, which collectively contain 14K hours of speech, and leverages LLMs al…

2024

Multi-Modal Retrieval For Large Language Model Based Speech Recognition

ACL 2024findings

Retrieval is a widely adopted approach for improving language models leveraging external information. As the field moves towards multi-modal large language models, it is important to extend the pure text based methods to incorporate other modalities in retrieval as well for applications across the w…

2023

Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers

ICML 2023poster

Streaming speech recognition architectures are employed for low-latency, real-time applications. Such architectures are often characterized by their causality. Causal architectures emit tokens at each frame, relying only on current and past signal, while non-causal models are exposed to a window of…

Cited by 9SourcePDFScholar