← Search

Gallil Maimon

3 accepted papers

2025

Slamming: Training a Speech Language Model on One GPU in a Day

ACL 2025finding

We introduce *Slam*, a recipe for training high-quality Speech Language Models (SLMs) on a single academic GPU in 24 hours. We do so through empirical analysis of model initialisation and architecture, synthetic training data, preference optimisation with synthetic data and tweaking all other compon…

2023

Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised Units

EMNLP 2023long findings

We introduce DISSC, a novel, lightweight method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. Unlike DISSC, most voice conversion (VC) methods focus primarily on timbre, and ignore people's unique speaking style (prosody). The proposed ap…

Cited by 0SourcecodeScholar