← Search

Laurent Mazaré

3 accepted papers

2026

Vision-Speech Models: Teaching Speech Models to Converse about Images

CVPR 2026

The recent successes of Vision-Language models raise the question of how to equivalently imbue a pretrained speech model with vision understanding, an important milestone towards building a multimodal speech model able to freely converse about images. Building such a conversational Vision-Speech mod

Cited by 0SourcecodeScholar
2025

Aligning Spoken Dialogue Models from User Interactions

ICML 2025poster

We propose a novel preference alignment framework for improving spoken dialogue models on real-time conversations from user interactions. Current preference learning methods primarily focus on text-based language models, and are not directly suited to the complexities of real-time speech interaction…

Cited by 0SourcePDFScholar
2025

High-Fidelity Simultaneous Speech-To-Speech Translation

ICML 2025poster

We introduce Hibiki, a decoder-only model for simultaneous speech translation. Hibiki leverages a multistream language model to synchronously process source and target speech, and jointly produces text and audio tokens to perform speech-to-text and speech-to-speech translation. We furthermore addres…