← Search

Zack Zukowski

3 accepted papers

2026

LOW-RESOURCE GUIDANCE FOR CONTROLLABLE LATENT AUDIO DIFFUSION

ICASSP 2026poster

Generative audio requires fine-grained controllable outputs, yet most existing methods require model retraining on specific controls or inference-time controls (\textit{e.g.}, guidance) that can also be computationally demanding. By examining the bottlenecks of existing guidance-based controls, in p…

Cited by 0SourcePDFScholar
2025

Scaling Transformers for Low-Bitrate High-Quality Speech Coding

ICLR 2025poster

The tokenization of audio with neural audio codec models is a vital part of modern AI pipelines for the generation or understanding of speech, alone or in a multimodal context. Traditionally such tokenization models have concentrated on low parameter-count architectures using only components with st…