← Search

Mingyue Niu

7 accepted papers

2026

PGCSPose: Physics-Constrained Generation and Causal Semantic Fusion for Robust In-Hand Pose Estimation

RA-L 2026

Accurate in-hand pose estimation is essential for dexterous robotic manipulation but remains fragile under severe visual occlusion (<inline-formula><tex-math notation="LaTeX">$>$</tex-math></inline-formula>50%) and intermittent tactile contact. Existing visuo-tactile fusion methods treat vision and

Cited by 0SourceScholar
2022

Automatic Depression Level Assessment from Speech By Long-Term Global Information Embedding

ICASSP 2022accepted

Depression is a serious mood disorder which brings negative effects on people's social activities. Therefore, growing attention has been paid to automatic depression assessment, especially from speech. However, most of the previous work uses hand-crafted features or deep neural network-based feature…

Cited by 0SourceScholar
2022

Automatic Respiratory Sound Classification Via Multi-Branch Temporal Convolutional Network

ICASSP 2022accepted

Automated classification of respiratory sounds has become an active research area in recent years. While recent studies have utilised deep learning methods to aid with respiratory sound classification, the performance is heavily influenced by the datasets available for respiratory sound classificati…

Cited by 0SourceScholar
2022

Csenet: Complex Squeeze-and-Excitation Network for Speech Depression Level Prediction

ICASSP 2022accepted

Automatic speech depression level prediction (SDLP) is a very challenging problem in affective computing. There are many studies that have acquired quite good performances for SDLP. However, most of the input speech features of these studies are based on the amplitude spectrogram, which loses the ph…

Cited by 0SourceScholar
2021

Multi-Scale and Multi-Region Facial Discriminative Representation for Automatic Depression Level Prediction

ICASSP 2021accepted

Physiological studies have shown that differences in facial activities between depressed patients and normal individuals are manifested in different local facial regions and the durations of these activities are not the same. But most previous works extract features from the entire facial region at…

Cited by 0SourceScholar
2020

Multimodal Transformer Fusion for Continuous Emotion Recognition

ICASSP 2020accepted

Multimodal fusion increases the performance of emotion recognition because of the complementarity of different modalities. Compared with decision level and feature level fusion, model level fusion makes better use of the advantages of deep neural networks. In this work, we utilize the Transformer mo…

Cited by 0SourceScholar
2019

Discriminative Video Representation with Temporal Order for Micro-expression Recognition

ICASSP 2019accepted

Micro-expression recognition is a challenging task due to its low intensity and short duration and how to extract the subtle facial changes is a key issue in this field. Although there are many methods attempt to cope with this problem, they are difficult to encode the temporal order of all frames i…

Cited by 0SourceScholar