2026
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
ICLR 2026poster
Recent efforts target spoken language models (SLMs) that not only listen but also speak for more natural human-LLM interaction. Joint text-speech modeling is a promising direction to achieve this. However, the effectiveness of recent speech tokens for joint modeling remains under-explored. To addres…