← Search

Marcely Zanon boito

4 accepted papers

2025

From Tower to Spire: Adding the Speech Modality to a Translation-Specialist LLM

EMNLP 2025

We introduce Spire, a speech-augmented language model (LM) capable of both translating and transcribing speech input from English into 10 other languages as well as translating text input in both language directions. Spire integrates the speech modality into an existing multilingual LM via speech di

2024

Multilingual Distilwhisper: Efficient Distillation of Multi-Task Speech Models Via Language-Specific Experts

ICASSP 2024accepted

Whisper is a multitask and multilingual speech model covering 99 languages. It yields commendable automatic speech recognition (ASR) results in a subset of its covered languages, but the model still underperforms on a non-negligible number of under-represented languages, a problem exacerbated in sma…

Cited by 0SourceScholar
2021

Task Agnostic and Task Specific Self-Supervised Learning from Speech with LeBenchmark

NeurIPS 2021poster

Self-Supervised Learning (SSL) has yielded remarkable improvements in many different domains including computer vision, natural language processing and speech processing by leveraging large amounts of unlabeled data. In the specific context of speech, however, and despite promising results, there ex…

Cited by 41SourceScholar