← Search

Ngoc-Quan Pham

7 accepted papers

2025

Improving Generalization with Flat Hilbert Bayesian Inference

ICML 2025poster

We introduce Flat Hilbert Bayesian Inference (FHBI), an algorithm designed to enhance generalization in Bayesian inference. Our approach involves an iterative two-step procedure with an adversarial functional perturbation step and a functional descent step within the reproducing kernel Hilbert space…

Cited by 0SourcePDFScholar
2025

Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS

ICASSP 2025accepted

Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciation issues, which can make it difficult for listeners to understand them. Hence,…

Cited by 0SourceScholar
2025

PIER: A Novel Metric for Evaluating What Matters in Code-Switching

ICASSP 2025accepted

Code-switching, the alternation of languages within a single discourse, presents a significant challenge for Automatic Speech Recognition. Despite the unique nature of the task, performance is commonly measured with established metrics such as Word-Error-Rate (WER). However, in this paper, we questi…

Cited by 0SourceScholar
2025

Promoting Ensemble Diversity with Interactive Bayesian Distributional Robustness for Fine-tuning Foundation Models

ICML 2025poster

We introduce Interactive Bayesian Distributional Robustness (IBDR), a novel Bayesian inference framework that allows modeling the interactions between particles, thereby enhancing ensemble quality through increased particle diversity. IBDR is grounded in a generalized theoretical framework that conn…

Cited by 0SourcePDFScholar
2024

DECM: Evaluating Bilingual ASR Performance on a Code-switching/mixing Benchmark

COLING 2024main

Automatic Speech Recognition has made significant progress, but challenges persist. Code-switched (CSW) Speech presents one such challenge, involving the mixing of multiple languages by a speaker. Even when multilingual ASR models are trained, each utterance on its own usually remains monolingual. W…

2023

SYNTACC : Synthesizing Multi-Accent Speech By Weight Factorization

ICASSP 2023accepted

Conventional multi-speaker text-to-speech synthesis (TTS) is known to be capable of synthesizing speech for multiple voices, yet it cannot generate speech in different accents. This limitation has motivated us to develop SYNTACC (Synthesizing speech with accents) which adapts conventional multi-spea…

Cited by 0SourceScholar