← Search

Shah Nawaz

6 accepted papers

2026

DISTILLATION-BASED LAYER DROPPING (DLD): EFFECTIVE END-TO-END FRAMEWORK FOR DYNAMIC SPEECH NETWORKS

ICASSP 2026poster

Edge devices operate in constrained and varying resource settings, requiring dynamic architectures that can adapt to limitations of the available resources. To meet such demands, layer dropping ($\mathcal{LD}$) approach is typically used to transform static models into dynamic ones by skipping parts…

Cited by 0SourcePDFScholar
2026

Linking Faces and Voices Across Languages: Insights from the FAME 2026 Challenge

ICASSP 2026poster

Over half of the world's population is bilingual and people often communicate under multilingual scenarios. The Face-Voice Association in Multilingual Environments (FAME) 2026 Challenge, held at ICASSP 2026, focuses on developing methods for face-voice association that are effective when the languag…

Cited by 0SourcePDFScholar
2026

Towards Effective Waste Segmentation for Automated Waste Recycling in Cluttered Background

ICML 2026poster

Rapid expansion of urban areas and population growth is causing an immense increase in waste production, which demands the need for efficient and automated waste management. In this scenario, automated waste recycling (AWR) that utilizes deep learning methods to separate the recyclable waste objects…

Cited by 0SourceScholar
2024

Frame-to-Utterance Convergence: A Spectra-Temporal Approach for Unified Spoofing Detection

ICASSP 2024accepted

Voice spoofing attacks pose a significant threat to automated speaker verification systems. Existing anti-spoofing methods often simulate specific attack types, such as synthetic or replay attacks. However, in real-world scenarios, the countermeasures are unaware of the generation schema of the atta…

Cited by 0SourceScholar
2023

Single-branch Network for Multimodal Training

ICASSP 2023accepted

With the rapid growth of social media platforms, users are sharing billions of multimedia posts containing audio, images, and text. Researchers have focused on building autonomous systems capable of processing such multimedia data to solve challenging multimodal tasks including cross-modal retrieval…

Cited by 0SourceScholar
2022

Fusion and Orthogonal Projection for Improved Face-Voice Association

ICASSP 2022accepted

We study the problem of learning association between face and voice. Prior works adopt pairwise or triplet loss formulations to learn an embedding space amenable for associated matching and verification tasks. Albeit showing some progress, such loss formulations are restrictive due to dependency on…

Cited by 0SourceScholar