← Search

Jiaxing Liu

6 accepted papers

2026

TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation

ICRA 2026poster

Vision-Language Navigation (VLN) presents a unique challenge for Large Vision-Language Models (VLMs) due to their inherent architectural mismatch: VLMs are primarily pretrained on static, disembodied vision-language tasks, which fundamentally clash with the dynamic, embodied, and spatially-structure…

2022

Domain-Invariant Feature Learning for Cross Corpus Speech Emotion Recognition

ICASSP 2022accepted

To deal with speech emotion recognition (SER) in real-life applications, researchers have to focus on cross corpus SER, where the feature distribution of source and target datasets are different. In this paper, we propose an efficient domain adversarial training method to cope with the non-affective…

Cited by 0SourceScholar
2022

Multi-Stage Graph Representation Learning for Dialogue-Level Speech Emotion Recognition

ICASSP 2022accepted

With the development of speech emotion recognition (SER), most of current research is utterance-level and cannot fit the need of actual scenarios. In this paper, we propose a novel strategy that focuses on capturing dialogue-level contextual information. On the basis of utterance-level representatio…

Cited by 0SourceScholar
2021

Domain-Adversarial Autoencoder with Attention Based Feature Level Fusion for Speech Emotion Recognition

ICASSP 2021accepted

Over the past two decades, although speech emotion recognition (SER) has garnered considerable attention, the problem of insufficient training data has been unresolved. A potential solution for this problem is to pre-train a model and transfer knowledge from large amounts of audio data. However, the…

Cited by 0SourceScholar
2021

Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation Fusion

ICASSP 2021accepted

Due to the more robust characteristics compared to unimodal, audio-video multimodal emotion recognition (MER) has attracted a lot of attention. The efficiency of representation fusion algorithm often determines the performance of MER. Although there are many fusion algorithms, information redundancy…

Cited by 0SourceScholar
2020

Speech Emotion Recognition with Local-Global Aware Deep Representation Learning

ICASSP 2020accepted

Convolutional neural network (CNN) based deep representation learning methods for speech emotion recognition (SER) have demonstrated great success. The basic design of CNN restricts the ability to model only local information well. Capsule network (CapsNet) can overcome the shortages of CNNs to capt…

Cited by 0SourceScholar