← Search

Ning Ma

17 accepted papers

2026

ESTIMATING RESPIRATORY EFFORT FROM NOCTURNAL BREATHING SOUNDS FOR OBSTRUCTIVE SLEEP APNOEA SCREENING

ICASSP 2026poster

Obstructive sleep apnoea (OSA) is a prevalent condition with significant health consequences, yet many patients remain undiagnosed due to the complexity and cost of over-night polysomnography. Acoustic-based screening provides a scalable alternative, yet performance is limited by environmental noise…

Cited by 0SourcePDFScholar
2026

Transfer Learning for Paediatric Sleep Apnoea Detection Using Physiology-Guided Acoustic Models

ICASSP 2026poster

Paediatric obstructive sleep apnoea (OSA) is clinically significant yet difficult to diagnose, as children poorly tolerate sensor-based polysomnography. Acoustic monitoring provides a non-invasive alternative for home-based OSA screening, but limited paediatric data hinders the development of robust…

Cited by 0SourcePDFScholar
2024

CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language Models

ACL 2024long

Multimodal large language models (MLLMs) have demonstrated promising results in a variety of tasks that combine vision and language. As these models become more integral to research and applications, conducting comprehensive evaluations of their capabilities has grown increasingly important. However…

Cited by 8SourcePDFScholar
2023

Partition Speeds Up Learning Implicit Neural Representations Based on Exponential-Increase Hypothesis

ICCV 2023poster

Implicit neural representations (INRs) aim to learn a continuous function (i.e., a neural network) to represent an image, where the input and output of the function are pixel coordinates and RGB/Gray values, respectively. However, images tend to consist of many objects whose colors are not perfectly…

Cited by 10PDFcodeScholar
2022

Auditory-Based Data Augmentation for end-to-end Automatic Speech Recognition

ICASSP 2022accepted

End-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human auditory inspired front-ends have also demonstrated improvement for automatic spee…

Cited by 0SourceScholar
2022

Learning Spatial-Preserved Skeleton Representations for Few-Shot Action Recognition

ECCV 2022poster

"Few-shot action recognition aims to recognize few-labeled novel action classes and attracts growing attentions due to practical significance. Human skeletons provide explainable and data-efficient representation for this problem by explicitly modeling spatial-temporal relations among skeleton joint…

2021

Exploiting Non-Negative Matrix Factorization for Binaural Sound Localization in the Presence of Directional Interference

ICASSP 2021accepted

This study presents a novel solution to the problem of binaural localization of a speaker in the presence of interfering directional noise and reverberation. Using a state-of-the-art binaural localization algorithm based on a deep neural network (DNN), we propose adding a source separation stage bas…

Cited by 0SourceScholar
2019

Deep Learning Features for Robust Detection of Acoustic Events in Sleep-disordered Breathing

ICASSP 2019accepted

Sleep-disordered breathing (SDB) is a serious and prevalent condition, and acoustic analysis via consumer devices (e.g. smartphones) offers a low-cost solution to screening for it. We present a novel approach for the acoustic identification of SDB sounds, such as snoring, using bottleneck features l…

Cited by 0SourceScholar
2019

End-to-end Binaural Sound Localisation from the Raw Waveform

ICASSP 2019accepted

A novel end-to-end binaural sound localisation approach is proposed which estimates the azimuth of a sound source directly from the waveform. Instead of employing hand-crafted features commonly employed for binaural sound localisation, such as the interaural time and level difference, our end-to-end…

Cited by 63SourceScholar
2017

Improving audio-visual speech recognition using deep neural networks with dynamic stream reliability estimates

ICASSP 2017accepted

Audio-visual speech recognition is a promising approach to tackling the problem of reduced recognition rates under adverse acoustic conditions. However, finding an optimal mechanism for combining multi-modal information remains a challenging task. Various methods are applicable for integrating acous…

Cited by 0SourceScholar
2016

Robust audiovisual speech recognition using noise-adaptive linear discriminant analysis

ICASSP 2016accepted

Automatic speech recognition (ASR) has become a widespread and convenient mode of human-machine interaction, but it is still not sufficiently reliable when used under highly noisy or reverberant conditions. One option for achieving far greater robustness is to include another modality that is unaffe…

Cited by 0SourceScholar
2015

A machine-hearing system exploiting head movements for binaural sound localisation in reverberant conditions

ICASSP 2015accepted

This paper is concerned with machine localisation of multiple active speech sources in reverberant environments using two (binaural) microphones. Such conditions typically present a problem for `classical' binaural models. Inspired by the human ability to utilise head movements, the current study in…

Cited by 0SourceScholar
2015

Robust localisation of multiple speakers exploiting head movements and multi-conditional training of binaural cues

ICASSP 2015accepted

This paper addresses the problem of localising multiple competing speakers in the presence of room reverberation, where sound sources can be positioned at any azimuth on the horizontal plane. To reduce the amount of front-back confusions which can occur due to the similarity of interaural time diffe…

Cited by 40SourceScholar