IN-SYNC: ADAPTATION OF SPEECH AWARE LARGE LANGUAGE MODELS FOR ASR WITH WORD LEVEL TIMESTAMP PREDICTIONS
Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is critical for applications such as captioning, media search, and…