← Search

Hung-yi Lee

108 accepted papers

2026

FULL-DUPLEX-BENCH V1.5: EVALUATING OVERLAP HANDLING FOR FULL-DUPLEX SPEECH MODELS

ICASSP 2026poster

Full-duplex spoken dialogue systems promise to transform human-machine interaction from a rigid, turn-based protocol into a fluid, natural conversation. However, the central challenge to realizing this vision, managing overlapping speech, remains critically under-evaluated. We introduce Full-Duplex-…

Cited by 0SourcePDFScholar
2026

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

ICLR 2026poster

Speech-to-Speech (S2S) models have shown promising dialogue capabilities, but their ability to handle paralinguistic cues—such as emotion, tone, and speaker attributes—and to respond appropriately in both content and style remains underexplored. Progress is further hindered by the scarcity of high-q…

Cited by 0SourceScholar
2026

STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models

ICLR 2026poster

Spoken Language Models (SLMs) are designed to take speech inputs and produce spoken responses. However, current SLMs lack the ability to perform an internal, unspoken thinking process before responding. In contrast, humans typically engage in complex mental reasoning internally, enabling them to com…

Cited by 0SourcecodeScholar
2026

TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling

ICLR 2026poster

Recent efforts target spoken language models (SLMs) that not only listen but also speak for more natural human-LLM interaction. Joint text-speech modeling is a promising direction to achieve this. However, the effectiveness of recent speech tokens for joint modeling remains under-explored. To addres…

Cited by 0SourcecodeScholar
2026

THE ICASSP 2026 HUMDIAL CHALLENGE: BENCHMARKING HUMAN-LIKE SPOKEN DIALOGUE SYSTEMS IN THE LLM ERA

ICASSP 2026poster

Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowing the gap between human-machine and human-human interactions. Achieving truly ``human-like'' communication necessitates…

Cited by 0SourcePDFScholar
2025

Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback

ACL 2025long

While textless Spoken Language Models (SLMs) have shown potential in end-to-end speech-to-speech modeling, they still lag behind text-based Large Language Models (LLMs) in terms of semantic coherence and relevance. This work introduces the Align-SLM framework, which leverages preference optimization…

2025

Audio-Aware Large Language Models as Judges for Speaking Styles

EMNLP 2025

Audio-aware large language models (ALLMs) can understand the textual and non-textual information in the audio input. In this paper, we explore using ALLMs as an automatic judge to assess the speaking styles of speeches. We use ALLM judges to evaluate the speeches generated by SLMs on two tasks: voic

2025

Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning

ICASSP 2025accepted

Recent advancements in large audio-language models (LALMs) have shown impressive capabilities in understanding and reasoning about audio and speech information. However, these models still face challenges, including hallucinating non-existent sound events, misidentifying the order of sound events, a…

Cited by 0SourceScholar
2025

Creativity in LLM-based Multi-Agent Systems: A Survey

EMNLP 2025

Large language model (LLM)-driven multi-agent systems (MAS) are transforming how humans and AIs collaboratively generate ideas and artifacts. While existing surveys provide comprehensive overviews of MAS infrastructures, they largely overlook the dimension of creativity , including how novel outputs

2025

Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

ICASSP 2025accepted

Recent end-to-end speech language models (SLMs) have expanded upon the capabilities of large language models (LLMs) by incorporating pre-trained speech models. However, these SLMs often undergo extensive speech instruction-tuning to bridge the gap between speech and text modalities. This requires si…

Cited by 0SourceScholar
2025

Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

ICLR 2025poster

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication…

2025

Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling

ICASSP 2025accepted

Multilingual Automatic Speech Recognition (ASR) aims to recognize and transcribe speech from multiple languages within a single system. By leveraging a vast amount of data and incorporating language tokens as prefixes to guide the recognition process, Whisper is one of the most advanced multilingual…

Cited by 0SourceScholar
2025

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction

ICML 2025poster

Autoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language Processing (NLP) at scale. Extending this paradigm to audio poses unique challenges due to its inherently continuous natu…

Cited by 0SourcePDFScholar
2025

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling

ICML 2025poster

Text-to-audio generation synthesizes realistic sounds or music given a natural language prompt. Diffusion-based frameworks, including the Tango and the AudioLDM series, represent the state-of-the-art in text-to-audio generation. Despite achieving high audio fidelity, they incur significant inference…

2025

Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection

ICASSP 2025accepted

Speech Emotion Recognition (SER) is a crucial component in developing general-purpose AI agents capable of natural human-computer interaction. However, building robust multilingual SER systems remains challenging due to the scarcity of labeled data in languages other than English and Chinese. In thi…

Cited by 0SourceScholar
2025

Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token Learning

NeurIPS 2025poster

Maintaining consistent model performance across domains is a fundamental challenge in machine learning. While recent work has explored using LLM-generated data for fine-tuning, its impact on cross-domain generalization remains poorly understood. This paper presents a systematic analysis revealing th…

Cited by 0SourceScholar
2025

Safeguard Fine-Tuned LLMs Through Pre- and Post-Tuning Model Merging

EMNLP 2025

Fine-tuning large language models (LLMs) for downstream tasks often leads to catastrophic forgetting, notably degrading the safety of originally aligned models. While some existing methods attempt to restore safety by incorporating additional safety data, the quality of such data typically falls sho

Cited by 0SourcePDFScholar
2025

Spectral-Aware Low-Rank Adaptation for Speaker Verification

ICASSP 2025accepted

Previous research has shown that the principal singular vectors of a pre-trained model’s weight matrices capture critical knowledge. In contrast, those associated with small singular values may contain noise or less reliable information. As a result, the LoRA-based parameter-efficient fine-tuning (P…

Cited by 0SourceScholar
2025

SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning

ICASSP 2025accepted

Instruction-based speech processing is becoming popular. Studies show that training with multiple tasks boosts performance, but collecting diverse, large-scale tasks and datasets is expensive. Thus, it is highly desirable to design a fundamental task that benefits other downstream tasks. This paper…

Cited by 7SourceScholar
2025

TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge

ACL 2025long

The LLM-as-a-judge paradigm uses large language models (LLMs) for automated text evaluation, assigning a score to the input based on scoring rubrics. Existing methods for fine-tuning LLM-as-a-judge use cross-entropy (CE) loss, which neglects the numeric nature of score prediction. Recent work addres…

2025

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey

EMNLP 2025

With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory tasks. While numerous benchmarks have emerged to assess LALMs’ performance, they rem

Cited by 0SourcePDFScholar
2025

Transferring Textual Preferences to Vision-Language Understanding through Model Merging

ACL 2025short

Large vision-language models (LVLMs) perform outstandingly across various multimodal tasks. However, their ability to evaluate generated content remains limited, and training vision-language reward models (VLRMs) with preference data is computationally expensive. This paper explores a training-free…

Cited by 0SourcePDFScholar
2024

AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models

ICASSP 2024accepted

Audio-visual representation learning aims to develop systems with human-like perception by utilizing correlation between auditory and visual information. However, current models often focus on a limited set of tasks, and generalization abilities of learned representations are unclear. To this end, w…

Cited by 0SourceScholar
2024

Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

ACL 2024long

In spoken dialogue, even if two current turns are the same sentence, their responses might still differ when they are spoken in different styles. The spoken styles, containing paralinguistic and prosodic information, mark the most significant difference between text and speech modality. When using t…

2024

Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages

ACL 2024long

Recently, the development of open-source large language models (LLMs) has advanced rapidly. Nevertheless, due to data constraints, the capabilities of most open-source LLMs are primarily focused on English. To address this issue, we introduce the concept of chat vector to equip pre-trained language…

2024

Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

ACL 2024findings

The sound codec’s dual roles in minimizing data transmission latency and serving as tokenizers underscore its critical importance.Recent years have witnessed significant developments in codec models.The ideal sound codec should preserve content, paralinguistics, speakers, and audio information.Howev…

2024

Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech

EMNLP 2024main

Deep Learning-based end-to-end Automatic Speech Recognition (ASR) has made significant strides but still struggles with performance on out-of-domain samples due to domain shifts in real-world scenarios. Test-Time Adaptation (TTA) methods address this issue by adapting models using test samples at in…

2024

DogeRM: Equipping Reward Models with Domain Knowledge through Model Merging

EMNLP 2024main

Reinforcement learning from human feedback (RLHF) is a popular strategy for aligning large language models (LLMs) with desired behaviors. Reward modeling is a crucial step in RLHF. However, collecting paired preference data for training reward models is often costly and time-consuming, especially fo…

2024

Dynamic-Superb: Towards a Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark For Speech

ICASSP 2024accepted

Text language models have shown remarkable zero-shot capability in generalizing to unseen tasks when provided with well-formulated instructions. However, existing studies in speech processing primarily focus on limited or specific tasks. Moreover, the lack of standardized benchmarks hinders a fair c…

Cited by 0SourceScholar
2024

I Need Help! Evaluating LLM’s Ability to Ask for Users’ Support: A Case Study on Text-to-SQL Generation

EMNLP 2024main

This study explores the proactive ability of LLMs to seek user support. We propose metrics to evaluate the trade-off between performance improvements and user burden, and investigate whether LLMs can determine when to request help under varying information availability. Our experiments show that wit…

2024

Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course

EMNLP 2024main

Using large language models (LLMs) for automatic evaluation has become an important evaluation method in NLP research. However, it is unclear whether these LLM-based evaluators can be effectively applied in real-world classrooms to assess student assignments. This empirical report shares how we use…

Cited by 8SourcePDFScholar
2024

Let Me Speak Freely? A Study On The Impact Of Format Restrictions On Large Language Model Performance.

EMNLP 2024industry

Structured generation, the process of producing content in standardized formats like JSON and XML, is widely utilized in real-world applications to extract key output information from large language models (LLMs).This study investigates whether such constraints on generation space impact LLMs’ abili…

Cited by 6SourcePDFScholar
2024

Merging Facts, Crafting Fallacies: Evaluating the Contradictory Nature of Aggregated Factual Claims in Long-Form Generations

ACL 2024findings

Long-form generations from large language models (LLMs) contain a mix of factual and non-factual claims, making evaluating factuality difficult.Prior works evaluate the factuality of a long paragraph by decomposing it into multiple facts, verifying those facts independently, and aggregating the resu…

2024

Meta-Diffu$B$: A Contextualized Sequence-to-Sequence Text Diffusion Model with Meta-Exploration

NeurIPS 2024poster

The diffusion model, a new generative modeling paradigm, has achieved significant success in generating images, audio, video, and text. It has been adapted for sequence-to-sequence text generation (Seq2Seq) through DiffuSeq, termed the S2S-Diffusion model. Existing S2S-Diffusion models predominantly…

2024

Multimodal Transformer Distillation for Audio-Visual Synchronization

ICASSP 2024accepted

Audio-visual synchronization aims to determine whether the mouth movements and speech in the video are synchronized. VocaLiST reaches state-of-the-art performance by incorporating multimodal Transformers to model audio-visual interact information. However, it requires high computing resources, makin…

Cited by 0SourceScholar
2024

On the Evaluation of Speech Foundation Models for Spoken Language Understanding

ACL 2024findings

The Spoken Language Understanding Evaluation (SLUE) suite of benchmark tasks was recently introduced to address the need for openresources and benchmarking of complex spoken language understanding (SLU) tasks, including both classification and sequence generation tasks, on natural speech. The benchm…

Cited by 6SourcePDFScholar
2024

Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue

ICASSP 2024accepted

Large Language Models (LLMs) have demonstrated superior abilities in tasks such as chatting, reasoning, and question-answering. However, standard LLMs may ignore crucial paralinguistic information, such as sentiment, emotion, and speaking style, which are essential for achieving natural, human-like…

Cited by 0SourceScholar
2024

REBORN: Reinforcement-Learned Boundary Segmentation with Iterative Training for Unsupervised ASR

NeurIPS 2024poster

Unsupervised automatic speech recognition (ASR) aims to learn the mapping between the speech signal and its corresponding textual transcription without the supervision of paired speech-text data. A word/phoneme in the speech signal is represented by a segment of speech signal with variable length an…

2024

Scalable Ensemble-Based Detection Method Against Adversarial Attacks For Speaker Verification

ICASSP 2024accepted

Automatic speaker verification (ASV) is highly susceptible to adversarial attacks. Purification modules are usually adopted as a pre-processing to mitigate adversarial noise. However, they are commonly implemented across diverse experimental settings, rendering direct comparisons challenging. This p…

Cited by 0SourceScholar
2024

SpeechDPR: End-To-End Spoken Passage Retrieval For Open-Domain Spoken Question Answering

ICASSP 2024accepted

Spoken Question Answering (SQA) is essential for machines to reply to user’s question by finding the answer span within a given spoken passage. SQA has been previously achieved without ASR to avoid recognition errors and Out-of-Vocabulary (OOV) problems. However, the real-world problem of Open-domai…

Cited by 0SourceScholar
2024

StreamBench: Towards Benchmarking Continuous Improvement of Language Agents

NeurIPS 2024poster

Recent works have shown that large language model (LLM) agents are able to improve themselves from experience, which is an important ability for continuous enhancement post-deployment. However, existing benchmarks primarily evaluate their innate capabilities and do not assess their ability to improv…

2024

Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition

EMNLP 2024main

Synthetic data is widely used in speech recognition due to the availability of text-to-speech models, which facilitate adapting models to previously unseen text domains. However, existing methods suffer in performance when they fine-tune an automatic speech recognition (ASR) model on synthetic data…

2024

Towards ASR Robust Spoken Language Understanding Through in-Context Learning with Word Confusion Networks

ICASSP 2024accepted

In the realm of spoken language understanding (SLU). numerous natural language understanding (NLU) methodologies have been adapted by supplying large language models (LLMs) with transcribed speech instead of conventional written text. In real-world scenarios, prior to input into an LLM. an automated…

Cited by 0SourceScholar
2024

Unveiling Narrative Reasoning Limits of Large Language Models with Trope in Movie Synopses

EMNLP 2024finding

Large language models (LLMs) equipped with chain-of-thoughts (CoT) prompting have shown significant multi-step reasoning capabilities in factual content like mathematics, commonsense, and logic. However, their performance in narrative reasoning, which demands greater abstraction capabilities, remain…

2024

Zero Resource Code-Switched Speech Benchmark Using Speech Utterance Pairs for Multiple Spoken Languages

ICASSP 2024accepted

We introduce a new zero resource code-switched speech bench-mark designed to assess the code-switching capabilities of self-supervised speech encoders directly. We showcase a baseline system of language modeling on discrete units to demonstrate how the code-switching abilities of speech encoders can…

Cited by 0SourceScholar
2023

Bridging Speech and Textual Pre-Trained Models With Unsupervised ASR

ICASSP 2023accepted

Spoken language understanding (SLU) is a task aiming to extract high-level semantics from spoken utterances. Previous works have investigated the use of speech self-supervised models and textual pre-trained models, which have shown reasonable improvements to various SLU tasks. However, because of th…

Cited by 0SourceScholar
2023

Cascading and Direct Approaches to Unsupervised Constituency Parsing on Spoken Sentences

ICASSP 2023accepted

Past work on unsupervised parsing is constrained to written form. In this paper, we present the first study on unsupervised spoken constituency parsing given unlabeled spoken sentences and unpaired textual data. The goal is to determine the spoken sentences’ hierarchical syntactic structure in the f…

Cited by 0SourceScholar
2023

Ensemble Knowledge Distillation of Self-Supervised Speech Models

ICASSP 2023accepted

Distilled self-supervised models have shown competitive performance and efficiency in recent years. However, there is a lack of experience in jointly distilling multiple self-supervised speech models. In our work, we performed Ensemble Knowledge Distillation (EKD) on various self-supervised speech m…

Cited by 0SourceScholar
2023

Euro: Espnet Unsupervised ASR Open-Source Toolkit

ICASSP 2023accepted

This paper describes the ESPnet Unsupervised ASR Open-source Toolkit (EURO), an end-to-end open-source toolkit for unsupervised automatic speech recognition (UASR). EURO adopts the state-of-the-art UASR learning method introduced by the Wav2vec-U, originally implemented at FAIRSEQ, which leverages s…

Cited by 0SourceScholar
2023

Hierarchical Programmatic Reinforcement Learning via Learning to Compose Programs

ICML 2023poster

Aiming to produce reinforcement learning (RL) policies that are human-interpretable and can generalize better to novel scenarios, Trivedi et al. (2021) present a method (LEAPS) that first learns a program embedding space to continuously parameterize diverse programs from a pre-generated program data…

Cited by 18SourcePDFScholar
2023

Introducing Semantics into Speech Encoders

ACL 2023long

Recent studies find existing self-supervised speech encoders contain primarily acoustic rather than semantic information. As a result, pipelined supervised automatic speech recognition (ASR) to large language model (LLM) systems achieve state-of-the-art results on semantic spoken language tasks by u…

Cited by 4SourcePDFScholar
2023

M-SpeechCLIP: Leveraging Large-Scale, Pre-Trained Models for Multilingual Speech to Image Retrieval

ICASSP 2023accepted

This work investigates the use of large-scale, English-only pre-trained models (CLIP and HuBERT) for multilingual image-speech retrieval. For non-English image-speech retrieval, we outperform the current state-of-the-art performance by a wide margin both when training separate models for each langua…

Cited by 0SourceScholar
2023

Personalized Lightweight Text-to-Speech: Voice Cloning with Adaptive Structured Pruning

ICASSP 2023accepted

Personalized TTS is an exciting and highly desired application that allows users to train their TTS voice using only a few recordings. However, TTS training typically requires many hours of recording and a large model, making it unsuitable for deployment on mobile devices. To overcome this limitatio…

Cited by 0SourceScholar
2023

SLUE Phase-2: A Benchmark Suite of Diverse Spoken Language Understanding Tasks

ACL 2023long

Spoken language understanding (SLU) tasks have been studied for many decades in the speech research community, but have not received as much attention as lower-level tasks like speech and speaker recognition. In this work, we introduce several new annotated SLU benchmark tasks based on freely availa…

2023

T5lephone: Bridging Speech and Text Self-Supervised Models for Spoken Language Understanding Via Phoneme Level T5

ICASSP 2023accepted

In Spoken language understanding (SLU), a natural solution is concatenating pre-trained speech models (e.g. HuBERT) and pretrained language models (PLM, e.g. T5). Most previous works use pre-trained language models with subword-based tokenization. However, the granularity of input units affects the…

Cited by 0SourceScholar
2022

AdapterBias: Parameter-efficient Token-dependent Representation Shift for Adapters in NLP Tasks

NAACL 2022findings

Transformer-based pre-trained models with millions of parameters require large storage. Recent approaches tackle this shortcoming by training adapters, but these approaches still require a relatively large number of parameters. In this study, AdapterBias, a surprisingly simple yet effective adapter…

2022

Adversarial Sample Detection for Speaker Verification by Neural Vocoders

ICASSP 2022accepted

Automatic speaker verification (ASV), one of the most important technology for biometric identification, has been widely adopted in security-critical applications. However, ASV is seriously vulnerable to recently emerged adversarial attacks, yet effective counter-measures against them are limited. I…

Cited by 0SourceScholar
2022

Analyzing The Robustness of Unsupervised Speech Recognition

ICASSP 2022accepted

Unsupervised speech recognition (unsupervised ASR) aims to learn the ASR system with non-parallel speech and text corpus only. Wav2vec-U [1] has shown promising results in unsupervised ASR by self-supervised speech representations coupled with Generative Adversarial Network (GAN) training, but the r…

Cited by 0SourceScholar
2022

Characterizing the Adversarial Vulnerability of Speech self-Supervised Learning

ICASSP 2022accepted

A leaderboard named Speech processing Universal PERformance Benchmark (SUPERB), which aims at benchmarking the performance of a shared self-supervised learning (SSL) speech model across various downstream speech tasks with minimal modification of architectures and a small amount of data, has fueled…

Cited by 0SourceScholar
2022

Distilhubert: Speech Representation Learning by Layer-Wise Distillation of Hidden-Unit Bert

ICASSP 2022accepted

Self-supervised speech representation learning methods like wav2vec 2.0 and Hidden-unit BERT (HuBERT) leverage unlabeled speech data for pre-training and offer good representations for numerous speech processing tasks. Despite the success of these methods, they require large memory and high pre-trai…

Cited by 0SourceScholar
2022

Don't Speak Too Fast: The Impact of Data Bias on Self-Supervised Speech Models

ICASSP 2022accepted

Self-supervised Speech Models (S3Ms) have been proven successful in many speech downstream tasks, like ASR. However, how pretraining data affects S3Ms’ downstream behavior remains an unexplored issue. In this paper, we study how pre-training data affects S3Ms by pre-training models on biased dataset…

Cited by 0SourceScholar
2022

On the Transferability of Pre-trained Language Models: A Study from Artificial Datasets

AAAI 2022technical

Pre-training language models (LMs) on large-scale unlabeled text data makes the model much easier to achieve exceptional downstream performance than their counterparts directly trained on the downstream tasks. In this work, we study what specific traits in the pre-training data, other than the sem…

2022

Partially Fake Audio Detection by Self-Attention-Based Fake Span Discovery

ICASSP 2022accepted

The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly implemented biometric identification models and can be harnessed by in-the-wild attackers for illegal uses. The ASVspoo…

Cited by 0SourceScholar
2022

S3PRL-VC: Open-Source Voice Conversion Framework with Self-Supervised Speech Representations

ICASSP 2022accepted

This paper introduces S3PRL-VC, an open-source voice conversion (VC) framework based on the S3PRL toolkit. In the context of recognition-synthesis VC, self-supervised speech representation (S3R) is valuable in its potential to replace the expensive supervised representation adopted by state-of-the-a…

Cited by 0SourceScholar
2022

SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities

ACL 2022long

Transfer learning has proven to be crucial in advancing the state of speech and natural language processing research in recent years. In speech, a model pre-trained by self-supervised learning transfers remarkably well on multiple tasks. However, the lack of a consistent evaluation methodology is li…

2022

XDBERT: Distilling Visual Information to BERT from Cross-Modal Systems to Improve Language Understanding

ACL 2022short

Transformer-based models are widely used in natural language understanding (NLU) tasks, and multimodal transformers have been effective in visual-language tasks. This study explores distilling visual information from pretrained multimodal transformers to pretrained language encoders. Our framework i…

Cited by 3SourcePDFScholar
2021

Adversarial Defense for Automatic Speaker Verification by Cascaded Self-Supervised Learning Models

ICASSP 2021accepted

Automatic speaker verification (ASV) is one of the core technologies in biometric identification. With the ubiquitous usage of ASV systems in safety-critical applications, more and more malicious attackers attempt to launch adversarial attacks at ASV systems. In the midst of the arms race between at…

Cited by 0SourceScholar
2021

Again-VC: A One-Shot Voice Conversion Using Activation Guidance and Adaptive Instance Normalization

ICASSP 2021accepted

Recently, voice conversion (VC) has been widely studied. Many VC systems use disentangle-based learning techniques to separate the speaker and the linguistic content information from a speech signal. Subsequently, they convert the voice by changing the speaker information to that of the target speak…

Cited by 0SourceScholar
2021

Fragmentvc: Any-To-Any Voice Conversion by End-To-End Extracting and Fusing Fine-Grained Voice Fragments with Attention

ICASSP 2021accepted

Any-to-any voice conversion aims to convert the voice from and to any speakers even unseen during training, which is much more challenging compared to one-to-one or many-to-many tasks, but much more attractive in real-world scenarios. In this paper we proposed FragmentVC, in which the latent phoneti…

Cited by 0SourceScholar
2021

Investigating on Incorporating Pretrained and Learnable Speaker Representations for Multi-Speaker Multi-Style Text-to-Speech

ICASSP 2021accepted

The few-shot multi-speaker multi-style voice cloning task is to synthesize utterances with voice and speaking style similar to a reference speaker given only a few reference samples. In this work, we investigate different speaker representations and proposed to integrate pretrained and learnable spe…

Cited by 0SourceScholar
2021

Is BERT a Cross-Disciplinary Knowledge Learner? A Surprising Finding of Pre-trained Models’ Transferability

EMNLP 2021finding

This paper investigates whether the power of the models pre-trained on text data, such as BERT, can be transferred to general token sequence classification applications. To verify pre-trained models’ transferability, we test the pre-trained models on text classification tasks with meanings of tokens…

2021

Put Chatbot into Its Interlocutor’s Shoes: New Framework to Learn Chatbot Responding with Intention

NAACL 2021long

Most chatbot literature that focuses on improving the fluency and coherence of a chatbot, is dedicated to making chatbots more human-like. However, very little work delves into what really separates humans from chatbots – humans intrinsically understand the effect their responses have on the interlo…

Cited by 7SourcePDFScholar
2021

Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model Pretraining

ICASSP 2021accepted

Much recent work on Spoken Language Understanding (SLU) is limited in at least one of three ways: models were trained on oracle text input and neglected ASR errors, models were trained to predict only intents without the slot values, or models were trained on a large amount of in-house data. In this…

Cited by 0SourceScholar
2020

Defense Against Adversarial Attacks on Spoofing Countermeasures of ASV

ICASSP 2020accepted

Various forefront countermeasure methods for automatic speaker verification (ASV) with considerable performance in anti-spoofing are proposed in the ASVspoof 2019 challenge. However, previous work has shown that countermeasure models are vulnerable to adversarial examples indistinguishable from natu…

Cited by 0SourceScholar
2020

Interrupted and Cascaded Permutation Invariant Training for Speech Separation

ICASSP 2020accepted

Permutation Invariant Training (PIT) has long been a stepping stone method for training speech separation model in handling the label ambiguity problem. With PIT selecting the minimum cost label assignments dynamically, very few studies considered the separation problem to be optimizing both the mod…

Cited by 0SourceScholar
2020

Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders

ICASSP 2020accepted

We present Mockingjay as a new speech representation learning approach, where bidirectional Transformer encoders are pre-trained on a large amount of unlabeled speech. Previous speech representation methods learn through conditioning on past frames and predicting information about future frames. Whe…

Cited by 0SourceScholar
2020

Self-Supervised Deep Learning for Fisheye Image Rectification

ICASSP 2020accepted

To rectify fisheye distortion from a single image, we advance self-supervised learning strategies and propose a unique deep learning model of Fisheye GAN (FE-GAN). Our FE-GAN learns pixel-level distortion flow from sets of fisheye distorted images and distortion-free ones (but not requiring such cor…

Cited by 0SourceScholar
2020

Sequence-to-Sequence Automatic Speech Recognition with Word Embedding Regularization and Fused Decoding

ICASSP 2020accepted

In this paper, we investigate the benefit that off-the-shelf word embedding can bring to the sequence-to-sequence (seq-to-seq) automatic speech recognition (ASR). We first introduced the word embedding regularization by maximizing the cosine similarity between a transformed decoder feature and the t…

Cited by 0SourceScholar
2020

TaylorGAN: Neighbor-Augmented Policy Update Towards Sample-Efficient Natural Language Generation

NeurIPS 2020poster

Score function-based natural language generation (NLG) approaches such as REINFORCE, in general, suffer from low sample efficiency and training instability problems. This is mainly due to the non-differentiable nature of the discrete space sampling and thus these methods have to treat the discrimina…

2020

Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning

ICASSP 2020accepted

In this paper we propose a Sequential Representation Quantization AutoEncoder (SeqRQ-AE) to learn from primarily unpaired audio data and produce sequences of representations very close to phoneme sequences of speech utterances. This is achieved by proper temporal segmentation to make the representat…

Cited by 0SourceScholar
2020

What Does a Network Layer Hear? Analyzing Hidden Representations of End-to-End ASR Through Speech Synthesis

ICASSP 2020accepted

End-to-end speech recognition systems have achieved competitive results compared to traditional systems. However, the complex transformations involved between layers given highly variable acoustic signals are hard to analyze. In this paper, we present our ASR probing model, which synthesizes speech…

Cited by 0SourceScholar
2019

Adversarial Learning of Label Dependency: A Novel Framework for Multi-class Classification

ICASSP 2019accepted

Recent work has shown that exploiting relations between labels improves the performance of multi-label classification. We propose a novel framework based on generative adversarial networks (GANs) to model label dependency. The discriminator learns to model label dependency by discriminating real and…

Cited by 0SourceScholar
2019

Adversarial Training of End-to-end Speech Recognition Using a Criticizing Language Model

ICASSP 2019accepted

In this paper we proposed a novel Adversarial Training (AT) approach for end-to-end speech recognition using a Criticizing Language Model (CLM). In this way the CLM and the automatic speech recognition (ASR) model can challenge and learn from each other iteratively to improve the performance. Since…

Cited by 0SourceScholar
2019

Mitigating the Impact of Speech Recognition Errors on Spoken Question Answering by Adversarial Domain Adaptation

ICASSP 2019accepted

Spoken question answering (SQA) is challenging due to complex reasoning on top of the spoken documents. The recent studies have also shown the catastrophic impact of automatic speech recognition (ASR) errors on SQA. Therefore, this work proposes to mitigate the ASR errors by aligning the mismatch be…

Cited by 0SourceScholar
2019

Towards Audio to Scene Image Synthesis Using Generative Adversarial Network

ICASSP 2019accepted

Humans can imagine a scene from a sound. We want machines to do so by using conditional generative adversarial networks (GANs). By applying the techniques including spectral norm, projection discriminator and auxiliary classifier, compared with naive conditional GAN, the model can generate images wi…

Cited by 0SourceScholar
2019

Towards End-to-end Speech-to-text Translation with Two-pass Decoding

ICASSP 2019accepted

Speech-to-text translation (ST) refers to transforming the audio in source language to the text in target language. Mainstream solutions for such tasks are to cascade automatic speech recognition with machine translation, for which the transcriptions of the source language are needed in training. En…

Cited by 0SourceScholar
2019

Using Deep-Q Network to Select Candidates from N-best Speech Recognition Hypotheses for Enhancing Dialogue State Tracking

ICASSP 2019accepted

Most state-of-the-art dialogue state tracking (DST) methods infer the dialogue state based on ground-truth transcriptions of utterances. In real-world situations, utterances are transcribed by automatic speech recognition (ASR) systems, which output the n-best candidate transcriptions (hypotheses).…

Cited by 0SourceScholar
2018

Domain Independent Key Term Extraction from Spoken Content Based on Context and Term Location Information in the Utterances

ICASSP 2018accepted

This paper proposes a domain independent approach for extracting key terms from spoken content based on context and term location information, or the sentence structures. Once it is trained with data of enough different domains, it is able to extract key terms in other unseen domains. This is obviou…

Cited by 0SourceScholar
2018

Language Transfer of Audio Word2Vec: Learning Audio Segment Representations Without Target Language Data

ICASSP 2018accepted

Audio Word2Vec offers vector representations of fixed dimensionality for variable-length audio segments using Sequence to-sequence Autoencoder (SA). These vector representations are shown to describe the sequential phonetic structures of the audio segments to a good degree, with real world applicati…

Cited by 0SourceScholar
2018

Scalable Sentiment for Sequence-to-Sequence Chatbot Response with Performance Analysis

ICASSP 2018accepted

Conventional seq2seq chatbot models only try to find the sentences with the highest probabilities conditioned on the input sequences, without considering the sentiment of the output sentences. Some research works trying to modify the sentiment of the output sequences were reported. In this paper, we…

Cited by 35SourceScholar
2018

Segmental Audio Word2Vec: Representing Utterances as Sequences of Vectors with Applications in Spoken Term Detection

ICASSP 2018accepted

While Word2Vec represents words (in text) as vectors carrying semantic information, audio Word2Vec was shown to be able to represent signal segments of spoken words as vectors carrying phonetic structure information. Audio Word2Vec can be trained in an unsupervised way from an unlabeled corpus, exce…

Cited by 0SourceScholar
2017

Personalized acoustic modeling by weakly supervised multi-task deep learning using acoustic tokens discovered from unlabeled data

ICASSP 2017accepted

It is well known that recognizers personalized to each user are much more effective than user-independent recognizers. With the popularity of smartphones today, although it is not difficult to collect a large set of audio data for each user, it is difficult to transcribe it. However, it is now possi…

Cited by 2SourceScholar
2017

Recurrent Neural Network based language modeling with controllable external Memory

ICASSP 2017accepted

It is crucial for language models to model long-term dependency in word sequences, which can be achieved to some good extent by recurrent neural network (RNN) based language models with long short-term memory (LSTM) units. To accurately model the sophisticated long-term information in human language…

Cited by 0SourceScholar