← Search

Benjamin Elizalde

12 accepted papers

2025

Audio Entailment: Assessing Deductive Reasoning for Audio Understanding

AAAI 2025technical

Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable performance in tasks including Text-to-Audio Retrieval, Captioning, and Question Answering. However, their ability to engage i…

2024

Natural Language Supervision For General-Purpose Audio Representations

ICASSP 2024accepted

Audio-Language models jointly learn multimodal text and audio representations that enable Zero-Shot inference. Models rely on the encoders to create powerful representations of the input and generalize to multiple tasks ranging from sounds, music, and speech. Although models have achieved remarkable…

Cited by 0SourceScholar
2024

Prompting Audios Using Acoustic Properties for Emotion Representation

ICASSP 2024accepted

Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propose the use of natural language descriptions (or prompts). In this work, we addres…

Cited by 0SourceScholar
2024

Training Audio Captioning Models without Audio

ICASSP 2024accepted

Automated Audio Captioning (AAC) is the task of generating natural language descriptions given an audio stream. A typical AAC system requires manually curated training data of audio segments and corresponding text caption annotations. The creation of these audio-caption pairs is costly, resulting in…

Cited by 0SourceScholar
2023

CLAP Learning Audio Concepts from Natural Language Supervision

ICASSP 2023accepted

Mainstream machine listening models are trained to learn audio concepts under the paradigm of one class label to many recordings focusing on one task. Learning under such restricted supervision limits the flexibility of models because they require labeled audio for training and can only predict the…

Cited by 0SourceScholar
2023

Multi-View Learning for Speech Emotion Recognition with Categorical Emotion, Categorical Sentiment, and Dimensional Scores

ICASSP 2023accepted

Psychological research has postulated that emotions and sentiment are correlated to dimensional scores of valence, arousal, and dominance. However, the literature of Speech Emotion Recognition focuses on independently predicting the three of them for a given speech audio. In this paper, we evaluate…

Cited by 0SourceScholar
2023

Pengi: An Audio Language Model for Audio Tasks

NeurIPS 2023poster

In the domain of audio processing, Transfer Learning has facilitated the rise of Self-Supervised Learning and Zero-Shot Learning techniques. These approaches have led to the development of versatile models capable of tackling a wide array of tasks, while delivering state-of-the-art performance. Howe…

2020

Multi-Label Sound Event Retrieval Using A Deep Learning-Based Siamese Structure With A Pairwise Presence Matrix

ICASSP 2020accepted

Realistic recordings of soundscapes often have multiple sound events co-occurring, such as car horns, engine and human voices. Sound event retrieval is a type of contentbased search aiming at finding audio samples, similar to an audio query based on their acoustic or semantic content. State of the a…

Cited by 0SourceScholar
2019

Cross Modal Audio Search and Retrieval with Joint Embeddings Based on Text and Audio

ICASSP 2019accepted

Existing audio search engines use one of two approaches: matching text-text or audio-audio pairs. In the former, text queries are matched to semantically similar words in an index of audio metadata to retrieve corresponding audio clips or segments, while in the latter, audio signals are directly use…

Cited by 65SourceScholar
2018

Acoustic Scene Classification Using Discrete Random Hashing for Laplacian Kernel Machines

ICASSP 2018accepted

State of the art acoustic scene classification techniques often employ features of large dimensionality, which are then used to train and perform inferences with kernel machines such as Support Vector Machines. However, the complexity of computing the non-linear kernel matrix for these methods incre…

Cited by 0SourceScholar
2018

Content-Based Representations of Audio Using Siamese Neural Networks

ICASSP 2018accepted

In this paper, we focus on the problem of content-based retrieval for audio, which aims to retrieve all semantically similar audio recordings for a given audio clip query. This problem is similar to the problem of query by example of audio, which aims to retrieve media samples from a database, which…

Cited by 0SourceScholar
2018

Framework for Evaluation of Sound Event Detection in Web Videos

ICASSP 2018accepted

The largest source of sound events is web videos. Most videos lack sound event labels at segment level, however, a significant number of them do respond to text queries, from a match found using metadata by search engines. In this paper we explore the extent to which a search query can be used as th…

Cited by 0SourceScholar