← Search

Zhao Ren

11 accepted papers

2025

Nonparametric Quantile Regression with ReLU-Activated Recurrent Neural Networks

NeurIPS 2025poster

This paper investigates nonparametric quantile regression using recurrent neural networks (RNNs) and sparse recurrent neural networks (SRNNs) to approximate the conditional quantile function, which is assumed to follow a compositional hierarchical interaction model. We show that RNN- and SRNN-based…

Cited by 0SourceScholar
2023

Cutting Through the Noise: An Empirical Comparison of Psycho-Acoustic and Envelope-based Features for Machinery Fault Detection

ICASSP 2023accepted

Acoustic-based fault detection has been one of the key instruments to monitor the health condition of mechanical parts. However, the background noise of an industrial environment may negatively influence the performance of fault detection. Limited attention has been paid to improving the robustness…

Cited by 0SourceScholar
2023

Fast Yet Effective Speech Emotion Recognition with Self-Distillation

ICASSP 2023accepted

Speech emotion recognition (SER) is the task of recognising humans’ emotional states from speech. SER is extremely prevalent in helping dialogue systems to truly understand our emotions and become a trustworthy human conversational partner. Due to the lengthy nature of speech, SER also suffers from…

Cited by 0SourceScholar
2023

Knowledge Transfer for on-Device Speech Emotion Recognition With Neural Structured Learning

ICASSP 2023accepted

Speech emotion recognition (SER) has been a popular research topic in human-computer interaction (HCI). As edge devices are rapidly springing up, applying SER to edge devices is promising for a huge number of HCI applications. Although deep learning has been investigated to improve the performance o…

Cited by 0SourceScholar
2022

Convoluational Transformer With Adaptive Position Embedding For Covid-19 Detection From Cough Sounds

ICASSP 2022accepted

Covid-19 has caused a huge health crisis worldwide in the past two years. Although an early detection of the virus through nucleic acid screening can considerably reduce its spread, the efficiency of this diagnostic process is limited by its complexity and costs. Hence, an effective and inexpensive…

Cited by 7SourceScholar
2020

Generating and Protecting Against Adversarial Attacks for Deep Speech-Based Emotion Recognition Models

ICASSP 2020accepted

The development of deep learning models for speech emotion recognition has become a popular area of research. Adversarially generated data can cause false predictions, and in an endeavor to ensure model robustness, defense methods against such attacks should be addressed. With this in mind, in this…

Cited by 0SourceScholar
2020

Latent Dynamic Factor Analysis of High-Dimensional Neural Recordings

NeurIPS 2020poster

High-dimensional neural recordings across multiple brain regions can be used to establish functional connectivity with good spatial and temporal resolution. We designed and implemented a novel method, Latent Dynamic Factor Analysis of High-dimensional time series (LDFA-H), which combines (a) a new a…

2019

Attention-based Atrous Convolutional Neural Networks: Visualisation and Understanding Perspectives of Acoustic Scenes

ICASSP 2019accepted

The goal of Acoustic Scene Classification (ASC) is to recognise the environment in which an audio waveform has been recorded. Recently, deep neural networks have been applied to ASC and have achieved state-of-the-art performance. However, few works have investigated how to visualise and understand w…

Cited by 0SourceScholar
2019

Implicit Fusion by Joint Audiovisual Training for Emotion Recognition in Mono Modality

ICASSP 2019accepted

Despite significant advances in emotion recognition from one individual modality, previous studies fail to take advantage of other modalities to train models in mono-modal scenarios. In this work, we propose a novel joint training model which implicitly fuses audio and visual information in the trai…

Cited by 0SourceScholar
2018

Towards Conditional Adversarial Training for Predicting Emotions from Speech

ICASSP 2018accepted

Motivated by the encouraging results recently obtained by generative adversarial networks in various image processing tasks, we propose a conditional adversarial training framework to predict dimensional representations of emotion, i. e., arousal and valence, from speech signals. The framework consi…

Cited by 0SourceScholar