← Search

K V Vijay Girish

1 accepted papers

2025

SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning

ACL 2025long

We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50M is built from publicly available speech corpora, which collectively contain 14K hours of speech, and leverages LLMs al…