← Search

Michel Olvera

3 accepted papers

2024

An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matching

NeurIPS 2024poster

Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In this work, we show that this ability can be re-purposed for audio captioning, where the joint image-language decoder can b…

Cited by 1SourcePDFScholar
2024

On The Choice of the Optimal Temporal Support for Audio Classification with Pre-Trained Embeddings

ICASSP 2024accepted

Current state-of-the-art audio analysis systems rely on pre-trained embedding models, often used off-the-shelf as (frozen) feature extractors. Choosing the best one for a set of tasks is the subject of many recent publications. However, one aspect often overlooked in these works is the influence of…

Cited by 0SourceScholar
2022

On The Impact of Normalization Strategies in Unsupervised Adversarial Domain Adaptation for Acoustic Scene Classification

ICASSP 2022accepted

Acoustic scene classification systems face performance degradation when trained and tested on data recorded by different devices. Unsupervised domain adaptation methods have been studied to reduce the impact of this mismatch. While they do not assume the availability of labels at test time, they oft…

Cited by 0SourceScholar