← Search

Karel Mundnich

8 accepted papers

2025

Speech Retrieval-Augmented Generation without Automatic Speech Recognition

ICASSP 2025accepted

One common approach for question answering over speech data is to first transcribe speech using automatic speech recognition (ASR) and then employ text-based retrieval-augmented generation (RAG) on the transcriptions. While this cascaded pipeline has proven effective in many practical settings, ASR…

Cited by 0SourceScholar
2025

Zero-resource Speech Translation and Recognition with LLMs

ICASSP 2025accepted

Despite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never…

Cited by 0SourceScholar
2024

SpeechGuard: Exploring the Adversarial Robustness of Multi-modal Large Language Models

ACL 2024findings

Integrated Speech and Large Language Models (SLMs) that can follow speech instructions and generate relevant text responses have gained popularity lately. However, the safety and robustness of these models remains largely unclear. In this work, we investigate the potential vulnerabilities of such in…

2020

Bringing in the Outliers: A Sparse Subspace Clustering Approach to Learn a Dictionary of Mouse Ultrasonic Vocalizations

ICASSP 2020accepted

Mice vocalize in the ultrasonic range during social interactions. These vocalizations are used in neuroscience and clinical studies to tap into complex behaviors and states. The analysis of these ultrasonic vocalizations (USVs) has been traditionally a manual process, which is prone to errors and hu…

Cited by 0SourceScholar
2020

The Role of Annotation Fusion Methods in the Study of Human-Reported Emotion Experience During Music Listening

ICASSP 2020accepted

Music is a universally-enjoyed art form, but listeners often respond to it in tremendously different ways. The same song can bring one person great joy and another deep sorrow. This paper focuses on modeling human music experience at the group level. In this scenario, human annotations serve an impo…

Cited by 0SourceScholar
2019

Bluetooth Based Indoor Localization Using Triplet Embeddings

ICASSP 2019accepted

We propose a novel algorithm for indoor localization using triplet embeddings through Bluetooth connectivity streams obtained in very noisy settings with irregular sampling schemes using environmental sensors distributed ad hoc inside buildings. We pose the problem as a matrix completion problem, wh…

Cited by 0SourceScholar
2018

A Novel Method for Human Bias Correction of Continuous- Time Annotations

ICASSP 2018accepted

Human annotations are of integral value in human behavior studies and in particular for the generation of ground truth for behavior prediction using various machine learning methods. These often subjective human annotations are especially required for studies involving measuring and predicting hidde…

Cited by 0SourceScholar