← Search

Hugo Flores García

2 accepted papers

2025

Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations

ICASSP 2025accepted

We present Sketch2Sound, a generative audio model capable of creating high-quality sounds from a set of interpretable time-varying control signals: loudness, brightness, and pitch, as well as text prompts. Sketch2Sound can synthesize arbitrary sounds from sonic imitations (i.e., a vocal imitation or…

Cited by 0SourceScholar
2025

Towards A Translative Model of Sperm Whale Vocalization

NeurIPS 2025poster

Sperm whales communicate in short sequences of clicks known as codas. We present WhAM (Whale Acoustics Model), the first transformer-based model capable of generating synthetic sperm whale codas from any audio prompt. WhAM is built by finetuning VampNet, a masked acoustic token model pretrained on m…

Cited by 1SourcecodeScholar