← Search

Syed Mohsen Naqvi

17 accepted papers

2025

A Frequency-aware Augmentation Network for Mental Disorders Assessment from Audio

ICASSP 2025accepted

Depression and Attention Deficit Hyperactivity Disorder (ADHD) stand out as the common mental health challenges today. In affective computing, speech signals serve as effective biomarkers for mental disorder assessment. Current research, relying on labor-intensive hand-crafted features or simplistic…

Cited by 0SourceScholar
2025

Efficient Long Speech Sequence Modelling for Time-Domain Depression Level Estimation

ICASSP 2025accepted

Depression significantly affects emotions, thoughts, and daily activities. Recent research indicates that speech signals contain vital cues about depression, sparking interest in audiobased deep-learning methods for estimating its severity. However, most methods rely on time-frequency representation…

Cited by 0SourceScholar
2024

Object Detection Oriented Privacy-Preserving Frame-Level Video Anomaly Detection

ICASSP 2024accepted

With the rapid development of intelligent surveillance, video anomaly detection has become a popular topic in related areas of artificial intelligence. In this work, the main focus is on those applications where the privacy of human targets is concerned, such as outdoor and indoor surveillance and s…

Cited by 8SourceScholar
2023

One-Shot Medical Action Recognition With A Cross-Attention Mechanism And Dynamic Time Warping

ICASSP 2023accepted

In this paper, we address the classification of medical actions with only one single sample by developing a novel one-shot learning framework which contains both cross-attention and dynamic time warping (DTW) modules. To be concrete, we firstly transform the raw skeleton sequence into the signal-lev…

Cited by 0SourceScholar
2022

A Two-Stream Information Fusion Approach to Abnormal Event Detection in Video

ICASSP 2022accepted

Human abnormal activity detection for automatic surveillance systems is to detect abnormal objects and human behaviours in videos. In this paper, we propose to explicitly address different kinds of abnormal events by developing a two-stream fusion approach that integrates both geometry and image tex…

Cited by 0SourceScholar
2020

Image Segmentation Based Privacy-Preserving Human Action Recognition for Anomaly Detection

ICASSP 2020accepted

Human Action Recognition and Anomaly Detection significantly improved automatic video analysis, assisted living, and video-based surveillance. The focus of this work is on those applications where privacy protection is required, such as surveillance and assisted living. RGB video data is the most co…

Cited by 0SourceScholar
2019

Privacy-preserving Online Human Behaviour Anomaly Detection Based on Body Movements and Objects Positions

ICASSP 2019accepted

Human behaviour anomaly detection is crucial for modern artifi-cial intelligence systems. However, privacy protection plays a great role in the realization. In this paper, an online privacy-preserving anomaly detector is presented. The proposed method is able to discriminate on human subject body mo…

Cited by 0SourceScholar
2018

3D-Hog Embedding Frameworks for Single and Multi-Viewpoints Action Recognition Based on Human Silhouettes

ICASSP 2018accepted

Given the high demand for automated systems for human action recognition, great efforts have been undertaken in recent decades to progress the field. In this paper, we present frameworks for single and multi-viewpoints action recognition based on Space-Time Volume (STV) of human silhouettes and 3D-H…

Cited by 0SourceScholar
2018

GM-PHD Filter Based Online Multiple Human Tracking Using Deep Discriminative Correlation Matching

ICASSP 2018accepted

In this paper, we propose deep discriminative correlation matching within the Gaussian Mixture Probability Hypothesis Density (GM-PHD) filter for online multiple human tracking. In this matching scheme, we mainly exploit the Convolutional Neural Network (CNN) based Discriminative Correlation Filter…

Cited by 0SourceScholar
2018

Geometric Information Based Monaural Speech Separation Using Deep Neural Network

ICASSP 2018accepted

The performance of deep neural network (DNN) based monaural speech separation methods is limited in reverberant and noisy room environments. In this paper, we propose a new DNN training target which incorporates geometric information describing the target speaker and microphone to improve the perfor…

Cited by 1SourceScholar
2017

Particle PHD filter based multi-target tracking using discriminative group-structured dictionary learning

ICASSP 2017accepted

Structured sparse representation has been recently found to achieve better efficiency and robustness in exploiting the target appearance model in tracking systems with both holistic and local information. Therefore, to better simultaneously discriminate multi-targets from their background, we propos…

Cited by 0SourceScholar
2017

Underdetermined source separation using time-frequency masks and an adaptive combined Gaussian-Student's t probabilistic model

ICASSP 2017accepted

Time-frequency (T-F) masking algorithms are focused at separating multiple sound sources from binaural reverberant speech mixtures. The statistical modelling of binaural cues i.e. interaural phase difference (IPD) and interaural level difference (ILD) is a significant aspect of such algorithms. In t…

Cited by 12SourceScholar
2016

Social force model aided robust particle PHD filter for multiple human tracking

ICASSP 2016accepted

In this paper, we propose a novel robust multiple human tracking approach based upon processing a video signal by utilizing a social force model to enhance the particle probability hypothesis density (PHD) filter. In traditional dynamic models, the states of targets are only predicted by their own h…

Cited by 0SourceScholar
2015

IVA algorithms using a multivariate Student's t source prior for speech source separation in real room environments

ICASSP 2015accepted

The independent vector analysis (IVA) algorithm employs a multivariate source prior to retain the dependency between different frequency bins of each source and thereby avoids the permutation problem that is inherent to blind source separation (BSS). In this paper, a multivariate Student's t distrib…

Cited by 12SourceScholar
2015

Real-time independent vector analysis with Student's t source prior for convolutive speech mixtures

ICASSP 2015accepted

A common approach to blind source separation is to use independent component analysis. However when dealing with realistic convolutive audio and speech mixtures, processing in the frequency domain at each frequency bin is required. As a result this introduces the permutation problem, inherent in ind…

Cited by 0SourceScholar
2015

Variational EM for clustering interaural phase cues in MESSL for blind source separation of speech

ICASSP 2015accepted

The model-based expectation maximization source separation and localization (MESSL) technique is a probabilistic time-frequency masking algorithm that achieves underdetermined blind source separation of speech sources. Using only two-channel recordings, MESSL clusters spectrogram points based on the…

Cited by 5SourceScholar