← Search

Dan Su

53 accepted papers

2025

Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

ACL 2025long

Recent English Common Crawl datasets like FineWeb-Edu and DCLM achieved significant benchmark gains via aggressive model-based filtering, but at the cost of removing 90% of data. This limits their suitability for long token horizon training, such as 15T tokens for Llama 3.1. In this paper, we show h…

2025

Nemotron-CLIMB: Clustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training

NeurIPS 2025spotlight

Pre-training datasets are typically collected from web content and lack inherent domain divisions. For instance, widely used datasets like Common Crawl do not include explicit domain labels, while manually curating labeled datasets such as The Pile is labor-intensive. Consequently, identifying an op…

Cited by 0SourceScholar
2024

DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis

ICASSP 2024accepted

This paper proposes an improved version of DurIAN-E (DurIAN-E 2), which is also a duration informed attention neural network for expressive and high-fidelity text-to-speech (TTS) synthesis. Similar with the DurIAN-E model, multiple stacked SwishRNN-based Transformer blocks are utilized as linguistic…

Cited by 0SourceScholar
2024

Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer

ACL 2024long

While recent advancements in speech language models have achieved significant progress, they face remarkable challenges in modeling the long acoustic sequences of neural audio codecs. In this paper, we introduce Generative Pre-trained Speech Transformer (GPST), a hierarchical transformer designed fo…

2024

MM-LLMs: Recent Advances in MultiModal Large Language Models

ACL 2024findings

In the past year, MultiModal Large Language Models (MM-LLMs) have undergone substantial advancements, augmenting off-the-shelf LLMs to support MM inputs or outputs via cost-effective training strategies. The resulting models not only preserve the inherent reasoning and decision-making capabilities o…

2024

Opine: Leveraging a Optimization-Inspired Deep Unfolding Method for Multi-Channel Speech Enhancement

ICASSP 2024accepted

Proximal gradient theory has demonstrated its superiority in the compressive sensing field for complex signal recovery. As an early trial in the speech front-end field, we propose OPINE, an optimization-inspired deep unfolding framework to simulate traditional iterative optimization process for mult…

Cited by 0SourceScholar
2024

Prompt-guided Precise Audio Editing with Diffusion Models

ICML 2024poster

Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio t…

Cited by 2SourcePDFScholar
2023

NusaCrowd: Open Source Initiative for Indonesian NLP Resources

ACL 2023findings

We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data loaders. The quality of the dataset…

2023

Trinet: Stabilizing Self-Supervised Learning From Complete or Slow Collapse

ICASSP 2023accepted

Self-supervised learning (SSL) models confront challenges of abrupt informational collapse or slow dimensional collapse. We propose TriNet, which introduces a novel triple-branch architecture for preventing collapse and stabilizing the pretraining. TriNet learns the SSL latent embedding space and in…

Cited by 0SourceScholar
2023

UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice Synthesis

AAAI 2023technical

Text-to-speech (TTS) and singing voice synthesis (SVS) aim at generating high-quality speaking and singing voice according to textual input and music scores, respectively. Unifying TTS and SVS into a single system is crucial to the applications requiring both of them. Existing methods usually suffer…

Cited by 10SourcePDFScholar
2022

BDDM: Bilateral Denoising Diffusion Models for Fast and High-Quality Speech Synthesis

ICLR 2022poster

Diffusion probabilistic models (DPMs) and their extensions have emerged as competitive generative models yet confront challenges of efficient sampling. We propose a new bilateral denoising diffusion model (BDDM) that parameterizes both the forward and reverse processes with a schedule network and a…

2022

Consistent Training and Decoding for End-to-End Speech Recognition Using Lattice-Free MMI

ICASSP 2022accepted

Recently, End-to-End (E2E) frameworks have achieved remarkable results on various Automatic Speech Recognition (ASR) tasks. However, Lattice-Free Maximum Mutual Information (LF-MMI), as one of the discriminative training criteria that show superior performance in hybrid ASR systems, is rarely adopte…

Cited by 0SourceScholar
2022

DP-DWA: Dual-Path Dynamic Weight Attention Network With Streaming Dfsmn-San For Automatic Speech Recognition

ICASSP 2022accepted

In multi-channel far-field automatic speech recognition (ASR) scenarios, distortion is introduced when the speech signal is processed by the front end, which damages the recognition performance for the ASR tasks. In this paper, we propose a dual-path network for the far-field acoustic model, which u…

Cited by 0SourceScholar
2022

Enhancing Speaking Styles in Conversational Text-to-Speech Synthesis with Graph-Based Multi-Modal Context Modeling

ICASSP 2022accepted

Comparing with traditional text-to-speech (TTS) systems, conversational TTS systems are required to synthesize speeches with proper speaking style confirming to the conversational context. However, state-of-the-art context modeling methods in conversational TTS only model the textual information in…

Cited by 0SourceScholar
2022

FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis

IJCAI 2022poster

Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hindered their applications to speech synthesis. This paper proposes FastDiff, a fast conditional diffusion model for high-qu…

2022

Multi-Channel Speaker Diarization Using Spatial Features for Meetings

ICASSP 2022accepted

Speaker identification for overlapped speech presents a great challenge for speaker diarization tasks in meeting scenarios. In order to overcome such challenges, several overlap-aware resegmentation methods based on deep learning have been integrated into speaker diarization systems. In this paper w…

Cited by 0SourceScholar
2022

Read before Generate! Faithful Long Form Question Answering with Machine Reading

ACL 2022findings

Long-form question answering (LFQA) aims to generate a paragraph-length answer for a given question. While current work on LFQA using large pre-trained model for generation are effective at producing fluent and somewhat relevant content, one primary challenge lies in how to generate a faithful answe…

2022

Referee: Towards Reference-Free Cross-Speaker Style Transfer with Low-Quality Data for Expressive Speech Synthesis

ICASSP 2022accepted

Cross-speaker style transfer (CSST) in text-to-speech (TTS) synthesis aims at transferring a speaking style to the synthesised speech in a target speaker’s voice. Most previous CSST approaches rely on expensive high-quality data carrying desired speaking style during training and require a re…

Cited by 0SourceScholar
2022

Simple Attention Module Based Speaker Verification with Iterative Noisy Label Detection

ICASSP 2022accepted

Recently, the attention mechanism such as squeeze-and-excitation module (SE) and convolutional block attention module (CBAM) has achieved great success in deep learning-based speaker verification system. This paper introduces an alternative effective yet simple one, i.e., simple attention module (Si…

Cited by 0SourceScholar
2022

The CUHK-Tencent Speaker Diarization System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge

ICASSP 2022accepted

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech recognition (ASR) tasks. In these meeting scenarios, the unce…

Cited by 0SourceScholar
2022

VCVTS: Multi-Speaker Video-to-Speech Synthesis Via Cross-Modal Knowledge Transfer from Voice Conversion

ICASSP 2022accepted

Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speech, while allowing flexible control of speaker identity, all in a single system. This paper proposes a novel multi-speake…

Cited by 0SourceScholar
2021

A Joint Training Framework of Multi-Look Separator and Speaker Embedding Extractor for Overlapped Speech

ICASSP 2021accepted

In multi-talker cases, overlapped speech degrades the speaker verification (SV) performance dramatically. To tackle this challenging problem, speech separation with multi-channel techniques can be adopted to extract each speaker’s signals to improve the SV performance. In this paper, a joint trainin…

Cited by 0SourceScholar
2021

Learned Transferable Architectures Can Surpass Hand-Designed Architectures for Large Scale Speech Recognition

ICASSP 2021accepted

In this paper, we explore the neural architecture search (NAS) for automatic speech recognition (ASR) systems. We conduct the architecture search on the small proxy dataset, and then evaluate the network, constructed from the searched architecture, on the large dataset. Specially, we propose a revis…

Cited by 0SourceScholar
2021

Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input

ICASSP 2021accepted

Non-autoregressive (NAR) transformer models have achieved significantly inference speedup but at the cost of inferior accuracy compared to autoregressive (AR) models in automatic speech recognition (ASR). Most of the NAR transformers take a fixed-length sequence filled with MASK tokens or a redundan…

Cited by 0SourceScholar
2021

Replay and Synthetic Speech Detection with Res2Net Architecture

ICASSP 2021accepted

Existing approaches for replay and synthetic speech detection still lack generalizability to unseen spoofing attacks. This work proposes to leverage a novel model structure, so-called Res2Net, to improve the anti-spoofing countermeasure’s generalizability. Res2Net mainly modifies the ResNet block to…

Cited by 0SourceScholar
2021

Sandglasset: A Light Multi-Granularity Self-Attentive Network for Time-Domain Speech Separation

ICASSP 2021accepted

One of the leading single-channel speech separation (SS) models is based on a TasNet with a dual-path segmentation technique, where the size of each segment remains unchanged throughout all layers. In contrast, our key finding is that multi-granularity features are essential for enhancing contextual…

Cited by 0SourceScholar
2021

Tune-In: Training Under Negative Environments with Interference for Attention Networks Simulating Cocktail Party Effect

AAAI 2021technical

We study the cocktail party problem and propose a novel attention network called Tune-In, abbreviated for training under negative environments with interference. It firstly learns two separate spaces of speaker-knowledge and speech-stimuli based on a shared feature space, where a new block structure…

2020

A Random Gossip BMUF Process for Neural Language Modeling

ICASSP 2020accepted

Neural network language model (NNLM) is an essential component of industrial ASR systems. One important challenge of training an NNLM is to leverage between scaling the learning process and handling big data. Conventional approaches such as block momentum provides a blockwise model update filtering…

Cited by 0SourceScholar
2020

Code-Switched Speech Synthesis Using Bilingual Phonetic Posteriorgram with Only Monolingual Corpora

ICASSP 2020accepted

Synthesizing fluent code-switched (CS) speech with consistent voice using only monolingual corpora is still a challenging task, since language alternation seldom occurs during training and the speaker identity is directly correlated with language. In this paper, we present a bilingual phonetic poste…

Cited by 0SourceScholar
2020

Dfsmn-San with Persistent Memory Model for Automatic Speech Recognition

ICASSP 2020accepted

Self-attention networks (SAN) have been introduced into automatic speech recognition (ASR) and achieved state-of-the-art performance owing to its superior ability in capturing long term dependency. One of the key ingredients is the self-attention mechanism which can be effectively performed on the w…

Cited by 0SourceScholar
2020

End-To-End Accent Conversion Without Using Native Utterances

ICASSP 2020accepted

Techniques for accent conversion (AC) aim to convert non-native to native accented speech. Conventional AC methods try to convert only the speaker identity of a native speaker's voice to that of the non-native accented target speaker, leaving the underlying content and pronunciations unchanged. This…

Cited by 0SourceScholar
2020

Enhancing End-to-End Multi-Channel Speech Separation Via Spatial Feature Learning

ICASSP 2020accepted

Hand-crafted spatial features (e.g., inter-channel phase difference, IPD) play a fundamental role in recent deep learning based multi-channel speech separation (MCSS) methods. However, these manually designed spatial features are hard to incorporate into the end-to-end optimized MCSS framework. In t…

Cited by 0SourceScholar
2020

Integration of Multi-Look Beamformers for Multi-Channel Keyword Spotting

ICASSP 2020accepted

Keyword spotting (KWS) is in great demand in smart devices in the era of Internet of Things. Albeit recent progresses, the performance of KWS, measured in false alarms and false rejects, may still degrade significantly under the far field and noisy conditions. In this paper, we propose integrating m…

Cited by 0SourceScholar
2020

Mixup-breakdown: A Consistency Training Method for Improving Generalization of Speech Separation Models

ICASSP 2020accepted

Deep-learning based speech separation models confront poor generalization problem that even the state-of-the-art models could abruptly fail when evaluating them in mismatch conditions. To address this problem, we propose an easy-to-implement yet effective consistency based semi-supervised learning (…

Cited by 0SourceScholar
2020

Multi-Level Deep Neural Network Adaptation for Speaker Verification Using MMD and Consistency Regularization

ICASSP 2020accepted

Adapting speaker verification (SV) systems to a new environment is a very challenging task. Current adaptation methods in SV mainly focus on the backend, i.e, adaptation is carried out after the speaker embeddings have been created. In this paper, we present a DNN-based adaptation method using maxim…

Cited by 0SourceScholar
2020

Speaker-Aware Target Speaker Enhancement by Jointly Learning with Speaker Embedding Extraction

ICASSP 2020accepted

Deep learning based speech separation approaches have received great interest, among which the recent speaker-aware speech enhancement methods are promising for solving difficulties such as arbitrary source permutation and unknown number of sources. In this paper, we propose a novel training framewo…

Cited by 0SourceScholar
2019

Boundary Discriminative Large Margin Cosine Loss for Text-independent Speaker Verification

ICASSP 2019accepted

Deep neural network based speaker embeddings have attracted much attention in text-independent speaker verification task. In addition to the network architecture, an appropriate design of the loss function is crucial for the deep discriminative embedding extractor. Inspired by the success of Large M…

Cited by 0SourceScholar
2019

Component Fusion: Learning Replaceable Language Model Component for End-to-end Speech Recognition System

ICASSP 2019accepted

Recently, attention-based end-to-end automatic speech recognition system (ASR) has shown promising results. One of the limitations of an attention-based ASR system is that its language model (LM) component has to be implicitly learned from transcribed speech data which prevents one from uti-lizing p…

Cited by 0SourceScholar
2019

DISR: Deep Infrared Spectral Restoration Algorithm for Robot Sensing and Intelligent Visual Tracking Systems

IROS 2019poster

Infrared imaging spectrometer (IRIS) often suffers from overlapped bands and random noises, which limit the precision of subsequent processing in robot vision sensing. To address this problem, we propose a novel Gabor transform-based infrared spectrum restoration method by successfully exploring the…

Cited by 12SourceScholar
2019

Investigating End-to-end Speech Recognition for Mandarin-english Code-switching

ICASSP 2019accepted

Code-switching is a common phenomenon in many multilingual communities and presents a challenge to automatic speech recognition (ASR). In this paper, three approaches are investigated to improve end-to-end speech recognition on Mandarin-English code-switching task. First, multi-task learning (MTL) i…

Cited by 0SourceScholar
2019

Joint Training of Complex Ratio Mask Based Beamformer and Acoustic Model for Noise Robust Asr

ICASSP 2019accepted

In this paper, we present a joint training framework between the multi-channel beamformer and the acoustic model for noise robust automatic speech recognition (ASR). The complex ratio mask (CRM), demonstrated to be more effective than the ideal ratio mask (IRM), is proposed to estimate the covarianc…

Cited by 0SourceScholar
2019

Learning Discriminative Features in Sequence Training without Requiring Framewise Labelled Data

ICASSP 2019accepted

In this work, we try to answer two questions: Can deeply learned features with discriminative power benefit an ASR system’s robustness to acoustic variability? And how to learn them without requiring framewise labelled sequence training data? As existing methods usually require knowing where the lab…

Cited by 0SourceScholar
2019

Multi-band PIT and Model Integration for Improved Multi-channel Speech Separation

ICASSP 2019accepted

The recent exploration of deep learning for supervised speech separation has significantly accelerated the progress on the multi-talker speech separation problem. Multi-channel extension has attracted much research attention due to the benefit of spatial information in far-field acoustic environment…

Cited by 0SourceScholar
2019

Quasi-fully Convolutional Neural Network with Variational Inference for Speech Synthesis

ICASSP 2019accepted

Recurrent neural networks, such as gated recurrent units (GRUs) and long short-term memory (LSTM), are widely used on acoustic modeling for speech synthesis. However, such sequential generating processes are not friendly to today’s massively parallel computing devices. We introduce a fully convoluti…

Cited by 0SourceScholar
2017

Development of precise mobile gaze tracking system based on online sparse Gaussian process regression and smooth-pursuit identification

ICRA 2017poster

In this paper, we aim to address two challenges in the implementation of mobile gaze tracking systems, i.e., the parallax error and the inflexible calibration procedure. Our proposed method mainly involves two steps and all the calibration process can be completed without needs to receive user's com…

Cited by 4SourceScholar
2017

M3Net: Multi-scale multi-path multi-modal fusion network and example application to RGB-D salient object detection

IROS 2017poster

Fusing RGB and depth data is compelling in boosting performance for various robotic and computer vision tasks. Typically, the streams of RGB and depth information are merged into a single fusion point in an early or late stage to generate combined features or decisions. The single fusion point also…

Cited by 14SourceScholar