← Search

Ming Lei

11 accepted papers

2026

Unleashing the Representational Power of Fourier Shapes for Attacking Infrared Object Detection

ICML 2026poster

Infrared object detection is crucial for perception in autonomous driving and surveillance but remains vulnerable to physical adversarial attacks. Unlike in the RGB domain, where attacks rely on color texture, infrared attacks must manipulate thermal signatures, making the geometry shape of heat-blo…

Cited by 0SourceScholar
2025

Improving Food Recognition with Retrieval-Augmented and Domain-Adaptive LVLMs

ICASSP 2025accepted

Food recognition is pivotal in enhancing intelligent food recommendation systems and nutritional management, contributing to balanced diets and overall health. Although Large Vision-Language Models (LVLMs) have demonstrated impressive performances across various domains, their performance on the foo…

Cited by 0SourceScholar
2022

Prosospeech: Enhancing Prosody with Quantized Vector Pre-Training in Text-To-Speech

ICASSP 2022accepted

Expressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling works have inevitable errors, which hurts the prosody modeling; 2) different attr…

Cited by 0SourceScholar
2019

Improving Audio-visual Speech Recognition Performance with Cross-modal Student-teacher Training

ICASSP 2019accepted

In this paper, we propose a cross-modal student-teacher learning framework to make a full use of externally abundant acoustic data in addition to a given task-specific audio-visual training database for improving speech recognition performance under the low signal-to-noise-ratio (SNR) and acoustic m…

Cited by 0SourceScholar
2019

Investigation of Modeling Units for Mandarin Speech Recognition Using Dfsmn-ctc-smbr

ICASSP 2019accepted

The choice of acoustic modeling units is critical to acoustic modeling in large vocabulary continuous speech recognition (LVCSR) tasks. The recent connectionist temporal classification (CTC) based acoustic models have more options for the choice of modeling units. In this work, we propose a DFSMN-CT…

Cited by 0SourceScholar
2019

Robust Audio-visual Speech Recognition Using Bimodal Dfsmn with Multi-condition Training and Dropout Regularization

ICASSP 2019accepted

Audio-visual speech recognition (AVSR) is thought to be one of the potential solutions for robust speech recognition, especially in noisy environments. Compared to audio only speech recognition, the major issues of AVSR include the lack of publicly available audio-visual corpora and the need of robu…

Cited by 0SourceScholar
2018

Deep Feed-Forward Sequential Memory Networks for Speech Synthesis

ICASSP 2018accepted

The Bidirectional LSTM (BLSTM) RNN based speech synthesis system is among the best parametric Text-to-Speech (TTS) systems in terms of the naturalness of generated speech, especially the naturalness in prosody. However, the model complexity and inference cost of BLSTM prevents its usage in many runt…

Cited by 0SourceScholar