← Search

Meng Yu

31 accepted papers

2026

Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement Learning

AAAI 2026technical

Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning utilizing rule-based rewards. However, the explicit reasoning process has not yet yielded su

Cited by 0SourcePDFScholar
2026

From Scale to Speed: Adaptive Test-Time Scaling for Image Editing

CVPR 2026

Image Chain-of-Thought (Image-CoT) is a test-time scaling paradigm that improves image generation by extending inference time. Most Image-CoT methods focus on text-to-image (T2I) generation. Unlike T2I generation, image editing is goal-directed: the solution space is constrained by the source image

Cited by 0SourceScholar
2025

BridgeVoC: Neural Vocoder with Schrödinger Bridge

IJCAI 2025

While previous diffusion-based neural vocoders typically follow a noise-to-data generation pipe-line, the linear-degradation prior of the mel-spectrogram is often neglected, resulting in limited generation quality. By revisiting the vocoding task and excavating its connection with the signal restora

Cited by 0SourcePDFScholar
2025

Neural Ambisonic Encoding For Multi-Speaker Scenarios Using A Circular Microphone Array

ICASSP 2025accepted

Spatial audio formats like Ambisonics are playback device layout-agnostic and well-suited for applications such as teleconferencing and virtual reality. Conventional Ambisonic encoding methods often rely on spherical microphone arrays for efficient sound field capture, which limits their flexibility…

Cited by 0SourceScholar
2025

ORA-NET: Enhancing Image Feature Matching through Oriented Overlapping Region Alignment

IROS 2025

Image feature matching is a fundamental task in computer vision. Existing local feature matching methods can establish robust correspondences between image pairs. However, these methods heavily rely on dense local image features, making them susceptible to significant perspective differences, charac

Cited by 0SourceScholar
2025

Open-RGBT: Open-Vocabulary RGB-T Zero-Shot Semantic Segmentation in Open-World Environments

ICRA 2025

Semantic segmentation is a critical technique for effective scene understanding. Traditional RGB-T semantic segmentation models often struggle to generalize across diverse scenarios due to their reliance on pretrained models and predefined categories. Recent advancements in Visual Language Models (V

Cited by 0SourcecodeScholar
2025

SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

ICASSP 2025accepted

In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the g…

Cited by 0SourceScholar
2025

TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models

IROS 2025

Reliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rely on low-level visual features and lack high-level textual information, which struggle with accurate segmentation when c

Cited by 1SourceScholar
2024

Advancing Acoustic Howling Suppression Through Recursive Training of Neural Networks

ICASSP 2024accepted

In this paper, we introduce a novel training framework designed to comprehensively address the acoustic howling issue by examining its fundamental formation process. This framework integrates a neural network (NN) module into the closed-loop system during training with signals generated recursively…

Cited by 0SourceScholar
2023

Deep Neural Mel-Subband Beamformer for in-Car Speech Separation

ICASSP 2023accepted

While current deep learning (DL)-based beamforming techniques have been proved effective in speech separation, they are often designed to process narrow-band (NB) frequencies independently which results in higher computational costs and inference times, making them unsuitable for real-world use. In…

Cited by 0SourceScholar
2022

Fast-Rir: Fast Neural Diffuse Room Impulse Response Generator

ICASSP 2022accepted

We present a neural-network-based fast diffuse room impulse response generator (FAST-RIR) for generating room impulse responses (RIRs) for a given acoustic environment. Our FAST-RIR takes rectangular room dimensions, listener and speaker positions, and reverberation time (T <inf xmlns:mml="http://ww…

Cited by 0SourceScholar
2022

Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization

ICASSP 2022accepted

Conversational bilingual speech encompasses three types of utterances: two purely monolingual types and one intra-sententially code-switched type. In this work, we propose a general framework to jointly model the likelihoods of the monolingual and code-switch sub-tasks that comprise bilingual speech…

Cited by 0SourceScholar
2022

Towards end-to-end Speaker Diarization with Generalized Neural Speaker Clustering

ICASSP 2022accepted

Speaker diarization consists of many components, e.g., front-end processing, speech activity detection (SAD), overlapped speech detection (OSD) and speaker segmentation/clustering. Conventionally, most of the involved components are separately developed and optimized. The resulting speaker diarizati…

Cited by 0SourceScholar
2021

A Joint Training Framework of Multi-Look Separator and Speaker Embedding Extractor for Overlapped Speech

ICASSP 2021accepted

In multi-talker cases, overlapped speech degrades the speaker verification (SV) performance dramatically. To tackle this challenging problem, speech separation with multi-channel techniques can be adopted to extract each speaker’s signals to improve the SV performance. In this paper, a joint trainin…

Cited by 0SourceScholar
2021

ADL-MVDR: All Deep Learning MVDR Beamformer for Target Speech Separation

ICASSP 2021accepted

Speech separation algorithms are often used to separate the target speech from other interfering sources. However, purely neural network based speech separation systems often cause nonlinear distortion that is harmful for automatic speech recognition (ASR) systems. The conventional mask-based minimu…

Cited by 0SourceScholar
2021

Directional ASR: A New Paradigm for E2E Multi-Speaker Speech Recognition with Source Localization

ICASSP 2021accepted

This paper proposes a new paradigm for handling far-field multi-speaker data in an end-to-end (E2E) neural network manner, called directional automatic speech recognition (D-ASR), which explicitly models source speaker locations. In D-ASR, the azimuth angle of the sources with respect to the microph…

Cited by 0SourceScholar
2021

Improving RNN Transducer with Target Speaker Extraction and Neural Uncertainty Estimation

ICASSP 2021accepted

Target-speaker speech recognition aims to recognize target-speaker speech from noisy environments with background noise and interfering speakers. This work presents a joint framework that combines time-domain target-speaker speech extraction and Recurrent Neural Network Transducer (RNN-T). To stabil…

Cited by 0SourceScholar
2021

Self-Supervised Text-Independent Speaker Verification Using Prototypical Momentum Contrastive Learning

ICASSP 2021accepted

In this study, we investigate self-supervised representation learning for speaker verification (SV). First, we examine a simple contrastive learning approach (SimCLR) with a momentum contrastive (MoCo) learning framework, where the MoCo speaker embedding system utilizes a queue to maintain a large s…

Cited by 0SourceScholar
2020

Enhancing End-to-End Multi-Channel Speech Separation Via Spatial Feature Learning

ICASSP 2020accepted

Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial features are hard to incorporate into the end-to-end optimized MCSS framework. In t…

Cited by 0SourceScholar
2020

Far-Field Location Guided Target Speech Extraction Using End-to-End Speech Recognition Objectives

ICASSP 2020accepted

Target speech extraction is a specific case of source separation where an auxiliary information like the location or some pre-saved anchor speech examples of the target speaker is used to resolve the permutation ambiguity. Traditionally such systems are optimized based on signal reconstruction objec…

Cited by 0SourceScholar
2020

Integration of Multi-Look Beamformers for Multi-Channel Keyword Spotting

ICASSP 2020accepted

Keyword spotting (KWS) is in great demand in smart devices in the era of Internet of Things. Albeit recent progresses, the performance of KWS, measured in false alarms and false rejects, may still degrade significantly under the far field and noisy conditions. In this paper, we propose integrating m…

Cited by 0SourceScholar
2020

Speaker-Aware Target Speaker Enhancement by Jointly Learning with Speaker Embedding Extraction

ICASSP 2020accepted

Deep learning based speech separation approaches have received great interest, among which the recent speaker-aware speech enhancement methods are promising for solving difficulties such as arbitrary source permutation and unknown number of sources. In this paper, we propose a novel training framewo…

Cited by 0SourceScholar
2019

Boundary Discriminative Large Margin Cosine Loss for Text-independent Speaker Verification

ICASSP 2019accepted

Deep neural network based speaker embeddings have attracted much attention in text-independent speaker verification task. In addition to the network architecture, an appropriate design of the loss function is crucial for the deep discriminative embedding extractor. Inspired by the success of Large M…

Cited by 0SourceScholar
2019

Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust Asr

ICASSP 2019accepted

In this paper, we present a joint training framework between the multi-channel beamformer and the acoustic model for noise robust automatic speech recognition (ASR). The complex ratio mask (CRM), demonstrated to be more effective than the ideal ratio mask (IRM), is proposed to estimate the covarianc…

Cited by 0SourceScholar
2019

Multi-band PIT and Model Integration for Improved Multi-channel Speech Separation

ICASSP 2019accepted

The recent exploration of deep learning for supervised speech separation has significantly accelerated the progress on the multi-talker speech separation problem. Multi-channel extension has attracted much research attention due to the benefit of spatial information in far-field acoustic environment…

Cited by 0SourceScholar
2019

Seq2Seq Attentional Siamese Neural Networks for Text-dependent Speaker Verification

ICASSP 2019accepted

In this paper, we present a Sequence-to-Sequence Attentional Siamese Neural Network (Seq2Seq-ASNN) that leverages temporal alignment information for end-to-end speaker verification. In prior works of speaker discriminative neural networks, utterance-level evaluation/enrollment speaker representation…

Cited by 0SourceScholar