← Search

Samarjit Das

8 accepted papers

2024

CLAP4Emo: ChatGPT-Assisted Speech Emotion Retrieval with Natural Language Supervision

ICASSP 2024accepted

Speech emotion retrieval is an important technique for large-scale and high-quality data collection. Conventional approach using ensemble of classification models might limit the retrieved emotion diversity and/or underperform in out-of-domain acoustic conditions. Natural language is diverse and agn…

Cited by 6SourceScholar
2022

Acoustic Imaging Aboard The International Space Station (ISS): Challenges and Preliminary Results

ICASSP 2022accepted

Design and execution of high fidelity acoustic sensing in complex environments poses a number of practical challenges, from accurately measuring the geometry of the setup and estimating the channel response, to time synchronization amongst the sources and receivers. When acoustic experiments are per…

Cited by 0SourceScholar
2022

Urban Sound & Sight: Dataset And Benchmark For Audio-Visual Urban Scene Understanding

ICASSP 2022accepted

Automatic audio-visual urban traffic understanding is a growing area of research with many potential applications of value to industry, academia, and the public sector. Yet, the lack of well-curated resources for training and evaluating models to research in this area hinders their development. To a…

Cited by 16SourceScholar
2021

Synthetic Aperture Acoustic Imaging with Deep Generative Model Based Source Distribution Prior

ICASSP 2021accepted

Acoustic imaging has a wide range of real-world applications such as machine health monitoring. Conventionally, large microphone arrays are utilized to achieve useful spatial resolution in the imaging process. The advent of location-aware autonomous mobile robotic platforms opens up unique opportuni…

Cited by 0SourceScholar
2018

A Light-Weight Multimodal Framework for Improved Environmental Audio Tagging

ICASSP 2018accepted

The lack of strong labels has severely limited the state-of-the-art fully supervised audio tagging systems to be scaled to larger dataset. Meanwhile, audio-visual learning models based on unlabeled videos have been successfully applied to audio tagging, but they are inevitably resource hungry and re…

Cited by 0SourceScholar
2018

Eventness: Object Detection on Spectrograms for Temporal Localization of Audio Events

ICASSP 2018accepted

In this paper, we introduce the concept of Eventness for audio event detection, which can, in part, be thought of as an analogue to Objectness from computer vision. The key observation behind the eventness concept is that audio events reveal themselves as 2-dimensional time-frequency patterns with s…

Cited by 0SourceScholar
2017

A comparison of Deep Learning methods for environmental sound detection

ICASSP 2017accepted

Environmental sound detection is a challenging application of machine learning because of the noisy nature of the signal, and the small amount of (labeled) data that is typically available. This work thus presents a comparison of several state-of-the-art Deep Learning models on the IEEE challenge on…

Cited by 0SourceScholar