← Search

Swapnil Bhosale

10 accepted papers

2026

MORE THAN A SHORTCUT: A HYPERBOLIC APPROACH TO EARLY-EXIT NETWORKS

ICASSP 2026poster

Deploying accurate event detection on resource-constrained devices is challenged by the trade-off between performance and computational cost. While Early-Exit (EE) networks offer a solution through adaptive computation, they often fail to enforce a coherent hierarchical structure, limiting the relia…

Cited by 0SourcePDFScholar
2025

Unsupervised Audio-Visual Segmentation with Modality Alignment

AAAI 2025technical

Audio-Visual Segmentation (AVS) aims to identify, at the pixel level, the object in a visual scene that produces a given sound. Current AVS methods rely on costly fine-grained annotations of mask-audio pairs, making them impractical for scalability. To address this, we propose the Modality Correspo…

2024

AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis

NeurIPS 2024poster

Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a condition for synthesizing binaural audio. However, in additio…

2024

Centrality-aware Product Retrieval and Ranking

EMNLP 2024industry

This paper addresses the challenge of improving user experience on e-commerce platforms by enhancing product ranking relevant to user’s search queries. Ambiguity and complexity of user queries often lead to a mismatch between user’s intent and retrieved product titles or documents. Recent approaches…

Cited by 0SourcePDFScholar
2024

DiffSED: Sound Event Detection with Denoising Diffusion

AAAI 2024technical

Sound Event Detection (SED) aims to predict the temporal boundaries of all the events of interest and their class labels, given an unconstrained audio sample. Taking either the split-and-classify (i.e., frame-level) strategy or the more principled event-level modeling approach, all existing methods…

2021

Deep Lung Auscultation Using Acoustic Biomarkers for Abnormal Respiratory Sound Event Detection

ICASSP 2021accepted

Lung Auscultation is a non-invasive process of distinguishing normal respiratory sounds from abnormal ones by analyzing the airflow along the respiratory tract. With developments in the Deep Learning (DL) techniques and wider access to anonymized medical data, automatic detection of specific sounds…

Cited by 0SourceScholar
2020

A Novel Approach for Intelligibility Assessment in Dysarthric Subjects

ICASSP 2020accepted

Dysarthria is a motor speech impairment caused by muscle weakness. Individuals, with this condition, are unable to control rapid movement of the velum leading to reduction in intelligibility, audibility, naturalness and efficiency of vocal communication. Systems that can assess intelligibility of dy…

Cited by 14SourceScholar
2020

Deep Encoded Linguistic and Acoustic Cues for Attention Based End to End Speech Emotion Recognition

ICASSP 2020accepted

An End-to-End model with convolutional layers and multi-head self attention mechanism is proposed for Speech Emotion Recognition (SER) task. As inputs, we propose to use both the deep encoded linguistic features that carry the language related context of emotion and the audio spectrogram that are re…

Cited by 0SourceScholar
2020

Improved Speaker Independent Dysarthria Intelligibility Classification Using Deepspeech Posteriors

ICASSP 2020accepted

Individuals with dysarthria are unable to control rapid movement of the velum leading to reduction in intelligibility, audibility, naturalness and efficiency of vocal communication. Automatic intelligibility assessment of dysarthric patients allows clinicians diagnose the impact of therapy and medic…

Cited by 0SourceScholar