← Search

Shuai Nie

14 accepted papers

2026

From Imitation to Discrimination: Toward a Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks

AAAI 2026technical

Reinforcement learning has emerged as a paradigm for post-training large language models, boosting their reasoning capabilities. Such approaches compute an advantage value for each sample, reflecting better or worse performance than expected, thereby yielding both positive and negative signals for t

Cited by 0SourcePDFScholar
2026

ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence On Mobile Devices

CVPR 2026

Multimodal large language models (MLLMs) have made significant progress in mobile agent development, yet their capabilities are predominantly confined to a reactive paradigm, where they merely execute explicit user commands. The emerging paradigm of proactive intelligence, where agents autonomously

Cited by 0SourcecodeScholar
2022

A Robust Deep Audio Splicing Detection Method via Singularity Detection Feature

ICASSP 2022accepted

There are many methods for detecting forged audio produced by conversion and synthesis. However, as a simpler method of forgery, splicing has not attracted widespread attention. Based on the characteristic that the tampering operation will cause singularities at high-frequency components, we propose…

Cited by 0SourceScholar
2022

ADD 2022: the first Audio Deep Synthesis Detection Challenge

ICASSP 2022accepted

Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three t…

Cited by 0SourceScholar
2021

Codebook Design for Dual-Polarized Ultra-Massive Mimo Communications at Millimeter Wave and Terahertz Bands

ICASSP 2021accepted

Intelligent surfaces at millimeter wave and terahertz band employ large antenna arrays based on advanced metamaterials to realize controllable wireless propagation. Such intelligent surfaces can achieve polarization tuning of electromagnetic waves, which allows polarization diversity under the ultra…

Cited by 0SourceScholar
2020

Beamforming in Intelligent Environments based on Ultra-Massive MIMO Platforms in Millimeter Wave and Terahertz Bands

ICASSP 2020accepted

Recent techniques utilizing reflectarrays or novel metamaterial-based surfaces to control wireless propagation environments have attracted great attention. While most solutions focus on sub-6 GHz wireless channel to assist in throughput enhancement, these intelligent surfaces have significant potent…

Cited by 0SourceScholar
2020

Mobility-Aware Beam Steering in Metasurface-Based Programmable Wireless Environments

ICASSP 2020accepted

Programmable wireless environments (PWEs) utilize electromagnetic metasurfaces to transform wireless propagation into a software-controlled resource. In this work we study the effects of user device mobility on the efficiency of PWEs. An analytical model is proposed, which describes the potential mi…

Cited by 0SourceScholar
2019

Adaptive Dereverberation Using Multi-channel Linear Prediction with Deficient Length Filter

ICASSP 2019accepted

In almost all adaptive dereverberation algorithms based on the multi-channel linear prediction (MCLP) model, it is assumed that the filter length can cover the reverberation time. However, in many practical situations, a deficient length filter, whose length is less than the reverberation time, is e…

Cited by 0SourceScholar
2019

Intelligent Environments Based on Ultra-massive Mimo Platforms for Wireless Communication in Millimeter Wave and Terahertz Bands

ICASSP 2019accepted

Millimeter-wave (30-300 GHz) and Terahertz-band communications (0.3-10 THz) are envisioned as key wireless technologies to satisfy the demand for Terabit-per-second (Tbps) links in the 5G and beyond eras. The very large available bandwidth in this ultra-broadband frequency range comes at the cost of…

Cited by 0SourceScholar
2019

Loss and Double-edge-triggered Detector for Robust Small-footprint Keyword Spotting

ICASSP 2019accepted

Keyword spotting (KWS) system constitutes a critical component of human-computer interfaces, which detects the specific keyword from a continuous stream of audio. The goal of KWS is providing a high detection accuracy at a low false alarm rate while having small memory and computation requirements.…

Cited by 17SourceScholar
2019

Sequence-To-Sequence Domain Adaptation Network for Robust Text Image Recognition

CVPR 2019poster

Domain adaptation has shown promising advances for alleviating domain shift problem. However, recent visual domain adaptation works usually focus on non-sequential object recognition with a global coarse alignment, which is inadequate to transfer effective knowledge for sequence-like text images wit…

Cited by 163PDFScholar
2018

Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training

ICASSP 2018accepted

In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a well-designed speech enhancement approach as the front-end of…

Cited by 0SourceScholar
2016

Exploiting spectro-temporal structures using NMF for DNN-based supervised speech separation

ICASSP 2016accepted

The targets of speech separation, whether ideal masks or magnitude spectrograms of interest, have prominent spectro-temporal structures. These characteristics are very worthy to be exploited for speech separation, however, they are usually ignored in previous works. In this paper, we use nonnegative…

Cited by 0SourceScholar
2015

A pairwise algorithm for pitch estimation and speech separation using deep stacking network

ICASSP 2015accepted

Pitch information is an important cue for speech separation. However, pitch estimation in noisy condition is also a task as challenging as speech separation. In this paper, we propose a supervised learning architecture which combines these two problems concisely. The proposed algorithm is based on d…

Cited by 0SourceScholar