← Search

Mingxing Xu

8 accepted papers

2026

GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs

ICLR 2026poster

Vision encoders are indispensable for allowing impressive performance of Multimodal Large Language Models (MLLMs) in vision–language tasks such as visual question answering and reasoning. However, existing vision encoders focus on global image representations but overlook fine-grained regional analy…

Cited by 0SourceScholar
2024

Enhancing Quantised End-to-End ASR Models Via Personalisation

ICASSP 2024accepted

Recent end-to-end automatic speech recognition (ASR) models have become increasingly larger, making them particularly challenging to be deployed on resource-constrained devices. Model quantisation is an effective solution that sometimes causes the word error rate (WER) to increase. In this paper, a…

Cited by 0SourceScholar
2021

Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language Understanding

EMNLP 2021main

Lack of training data presents a grand challenge to scaling out spoken language understanding (SLU) to low-resource languages. Although various data augmentation approaches have been proposed to synthesize training data in low-resource target languages, the augmented data sets are often noisy, and t…

2017

Learning cross-lingual knowledge with multilingual BLSTM for emphasis detection with limited training data

ICASSP 2017accepted

Bidirectional long short-term memory (BLSTM) recurrent neural network (RNN) has achieved state-of-the-art performance in many sequence processing problems given its capability in capturing contextual information. However, for languages with limited amount of training data, it is still difficult to o…

Cited by 0SourceScholar
2017

Speaker segmentation using deep speaker vectors for fast speaker change scenarios

ICASSP 2017accepted

A novel speaker segmentation approach based on deep neural network is proposed and investigated. This approach uses deep speaker vectors (d-vectors) to represent speaker characteristics and to find speaker change points. The d-vector is a kind of frame-level speaker discriminative feature, whose dis…

Cited by 0SourceScholar
2016

A deep bidirectional long short-term memory based multi-scale approach for music dynamic emotion prediction

ICASSP 2016accepted

Music Dynamic Emotion Prediction is a challenging and significant task. In this paper, We adopt the dimensional valence-arousal (V-A) emotion model to represent the dynamic emotion in music. Considering the high context correlation among the music feature sequence and the advantage of Bidirectional…

Cited by 0SourceScholar
2016

Question detection from acoustic features using recurrent neural network with gated recurrent unit

ICASSP 2016accepted

Question detection is of importance for many speech applications. Only parts of the speech utterances can provide useful clues for question detection. Previous work of question detection using acoustic features in Mandarin conversation is weak in capturing such proper time context information, which…

Cited by 0SourceScholar
2016

SVR based double-scale regression for dynamic emotion prediction in music

ICASSP 2016accepted

Dynamic music emotion prediction is to recognize the continuous emotion contained in music, and has various applications. In recent years, dynamic music emotion recognition is widely studied, while the inside structure of the emotion in music remains unclear. We conduct a data observation based on t…

Cited by 0SourceScholar