← Search

Ngoc Thang Vu

25 accepted papers

2026

HOW TO LABEL RESYNTHESIZED AUDIO: THE DUAL ROLE OF NEURAL AUDIO CODECS IN AUDIO DEEPFAKE DETECTION

ICASSP 2026poster

Since Text-to-Speech systems typically don't produce waveforms directly, recent spoof detection studies use resynthesized waveforms from vocoders and neural audio codecs to simulate an attacker. Unlike vocoders, which are specifically designed for speech synthesis, neural audio codecs were originall…

Cited by 0SourcePDFScholar
2025

A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions

COLING 2025main

Language in the Arab world presents a complex diglossic and multilingual setting, involving the use of Modern Standard Arabic, various dialects and sub-dialects, as well as multiple European languages. This diverse linguistic landscape has given rise to code-switching, both within Arabic varieties a…

Cited by 1SourcePDFScholar
2025

Discrete Subgraph Sampling for Interpretable Graph based Visual Question Answering

COLING 2025main

Explainable artificial intelligence (XAI) aims to make machine learning models more transparent. While many approaches focus on generating explanations post-hoc, interpretable approaches, which generate the explanations intrinsically alongside the predictions, are relatively rare. In this work, we i…

2025

High-Resolution Speech Restoration with Latent Diffusion Model

ICASSP 2025accepted

Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstruction and high-frequency harmonics, leading to breathing and gasping artifacts that…

Cited by 0SourceScholar
2025

It’s What You Say and How You Say It: Investigating the Effect of Linguistic vs. Behavioral Adaptation in Task-Oriented Chatbots

COLING 2025main

Given the conflicting expectations users have for how a dialog agent should sound and behave, there is no one-size-fits-all option for dialog system design. Therefore, adaptation is critical to ensure successful and enjoyable interactions. However, it is not yet clear what the effects of behavioral…

Cited by 0SourcePDFScholar
2025

What Affects the Performance of Fake Audio Detection? Analyzing Factors in a Continual Learning Setting

ICASSP 2025accepted

The increasing sophistication of deepfake audio generation technologies makes it important to develop robust fake audio detection systems that can adapt over time. This study examines how various factors impact the performance of detection systems in a continual learning setting. We focus on factors…

Cited by 0SourceScholar
2024

Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource Settings

COLING 2024main

Attribution scores indicate the importance of different input parts and can, thus, explain model behaviour. Currently, prompt-based models are gaining popularity, i.a., due to their easier adaptability in low-resource settings. However, the quality of attribution scores extracted from prompt-based m…

Cited by 2SourcePDFScholar
2024

Intrinsic Subgraph Generation for Interpretable Graph Based Visual Question Answering

COLING 2024main

The large success of deep learning based methods in Visual Question Answering (VQA) has concurrently increased the demand for explainable methods. Most methods in Explainable Artificial Intelligence (XAI) focus on generating post-hoc explanations rather than taking an intrinsic approach, the latter…

2024

Prompting-based Synthetic Data Generation for Few-Shot Question Answering

COLING 2024main

Although language models (LMs) have boosted the performance of Question Answering, they still need plenty of data. Data annotation, in contrast, is a time-consuming process. This especially applies to Question Answering, where possibly large documents have to be parsed and annotated with questions a…

2023

Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers

ACL 2023long

In recent years machine translation has become very successful for high-resource language pairs. This has also sparked new interest in research on the automatic translation of low-resource languages, including Indigenous languages. However, the latter are deeply related to the ethnic and cultural gr…

Cited by 28SourcePDFScholar
2023

Neighboring Words Affect Human Interpretation of Saliency Explanations

ACL 2023findings

Word-level saliency explanations (“heat maps over words”) are often used to communicate feature-attribution in text-based models. Recent studies found that superficial factors such as word length can distort human interpretation of the communicated saliency scores. We conduct a user study to investi…

2023

Prosody Is Not Identity: A Speaker Anonymization Approach Using Prosody Cloning

ICASSP 2023accepted

Prosody is closely linked to the identity of a speaker, leading to individual pitch and intonation patterns. Therefore, it is challenging in speaker anonymization to generate speech utterances that both keep the original audio’s main prosodic structure and preserve the speaker’s privacy. In this pap…

Cited by 0SourceScholar
2023

Regularisation for Efficient Softmax Parameter Generation in Low-Resource Text Classifiers

IJCAI 2023poster

Meta-learning has made tremendous progress in recent years and was demonstrated to be particularly suitable in low-resource settings where training data is very limited. However, meta-learning models still require large amounts of training tasks to achieve good generalisation. Since labelled trainin…

Cited by 0SourcePDFScholar
2022

ESPnet-SLU: Advancing Spoken Language Understanding Through ESPnet

ICASSP 2022accepted

As Automatic Speech Processing (ASR) systems are getting better, there is an increasing interest of using the ASR output to do downstream Natural Language Processing (NLP) tasks. However, there are few open source toolkits that can be used to generate reproducible results on different Spoken Languag…

Cited by 0SourceScholar
2021

Beyond Accuracy: A Consolidated Tool for Visual Question Answering Benchmarking

EMNLP 2021system demonstrations

On the way towards general Visual Question Answering (VQA) systems that are able to answer arbitrary questions, the need arises for evaluation beyond single-metric leaderboards for specific datasets. To this end, we propose a browser-based benchmarking tool for researchers and challenge organizers,…

2020

Fine-tuning BERT for Low-Resource Natural Language Understanding via Active Learning

COLING 2020main

Recently, leveraging pre-trained Transformer based language models in down stream, task specific models has advanced state of the art results in natural language understanding tasks. However, only a little research has explored the suitability of this approach in low resource settings with less than…

Cited by 66SourcePDFScholar
2020

OH, JEEZ! or UH-HUH? A Listener-Aware Backchannel Predictor on ASR Transcriptions

ICASSP 2020accepted

This paper presents our latest investigation on modeling backchannel in conversations. Motivated by a proactive backchanneling theory, we aim at developing a system which acts as a proactive listener by inserting backchannels, such as continuers and assessment, to influence speakers. Our model takes…

Cited by 0SourceScholar
2019

Context-aware Neural-based Dialog Act Classification on Automatically Generated Transcriptions

ICASSP 2019accepted

This paper presents our latest investigations on dialog act (DA) classification on automatically generated transcriptions. We propose a novel approach that combines convolutional neural networks (CNNs) and conditional random fields (CRFs) for context modeling in DA classification. We explore the imp…

Cited by 0SourceScholar
2019

Improving Speech Emotion Recognition with Unsupervised Representation Learning on Unlabeled Speech

ICASSP 2019accepted

In this paper we present our findings on how representation learning on large unlabeled speech corpora can be beneficially utilized for speech emotion recognition (SER). Prior work on representation learning for SER mostly focused on the relatively small emotional speech datasets without making use…

Cited by 0SourceScholar
2016

Bi-directional recurrent neural network with ranking loss for spoken language understanding

ICASSP 2016accepted

This paper presents our latest investigation of recurrent neural networks for the slot filling task of spoken language understanding. We implement a bi-directional Elman-type recurrent neural network which takes the information not only from the past but also from the future context to predict the s…

Cited by 0SourceScholar