← Search

Man-Wai Mak

22 accepted papers

2025

Denoising Student Features with Diffusion Models for Knowledge Distillation in Speaker Verification

ICASSP 2025accepted

In recent years, there has been a surge in the use of a pre-trained speech model as a feature extractor for speaker verification (SV). To reduce model complexity, researchers transfer knowledge from a pre-trained model to a lightweight student model, enabling the latter to reach a performance level…

Cited by 0SourceScholar
2025

Grouped Knowledge Distillation with Adaptive Logit Softening for Speaker Recognition

ICASSP 2025accepted

Recent works suggest that decoupling the information of non-target speakers from that of the target speaker in knowledge distillation (KD) and subsequently emphasizing the former can lead to significant performance improvement. However, a well-trained teacher model typically produces almost zero non…

Cited by 0SourceScholar
2025

Spectral-Aware Low-Rank Adaptation for Speaker Verification

ICASSP 2025accepted

Previous research has shown that the principal singular vectors of a pre-trained model’s weight matrices capture critical knowledge. In contrast, those associated with small singular values may contain noise or less reliable information. As a result, the LoRA-based parameter-efficient fine-tuning (P…

Cited by 0SourceScholar
2025

TrInk: Ink Generation with Transformer Network

EMNLP 2025

In this paper, we propose TrInk, a Transformer-based model for ink generation, which effectively captures global dependencies. To better facilitate the alignment between the input text and generated stroke points, we introduce scaled positional embeddings and a Gaussian memory mask in the cross-atte

2024

Asymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker Verification

ICASSP 2024accepted

Contrastive self-supervised learning (CSL) for speaker verification (SV) has drawn increasing interest recently due to its ability to exploit unlabeled data. Performing data augmentation on raw waveforms, such as adding noise or reverberation, plays a pivotal role in achieving promising results in S…

Cited by 7SourceScholar
2024

Dual Parameter-Efficient Fine-Tuning for Speaker Representation Via Speaker Prompt Tuning and Adapters

ICASSP 2024accepted

Fine-tuning a pre-trained Transformer model (PTM) for speech applications in a parameter-efficient manner offers the dual benefits of reducing memory and leveraging the rich feature representations in massive unlabeled datasets. However, existing parameter-efficient fine-tuning approaches either ada…

Cited by 0SourceScholar
2024

Promoting Independence of Depression and Speaker Features for Speaker Disentanglement in Speech-Based Depression Detection

ICASSP 2024accepted

Recent studies have demonstrated the effectiveness of speaker disentanglement in mitigating the interference caused by speaker features in speech-based depression detection. However, the inherent entanglement between depression features and speaker features poses challenges to depression detection.…

Cited by 0SourceScholar
2023

Discriminative Speaker Representation Via Contrastive Learning with Class-Aware Attention in Angular Space

ICASSP 2023accepted

The challenges in applying contrastive learning to speaker verification (SV) are that the softmax-based contrastive loss lacks discriminative power and that the hard negative pairs can easily influence learning. To overcome the first challenge, we propose a contrastive learning SV framework incorpor…

Cited by 0SourceScholar
2023

Feature Selection and Text Embedding for Detecting Dementia from Spontaneous Cantonese

ICASSP 2023accepted

Dementia is a severe cognitive impairment that affects the health of older adults and creates a burden on their families and caretakers. This paper analyzes diverse hand-crafted features extracted from spoken languages and selects the most discriminative ones for dementia detection. Recently, the pe…

Cited by 0SourceScholar
2023

Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech Representations

ICML 2023poster

Self-supervised learning (SSL) speech models such as wav2vec and HuBERT have demonstrated state-of-the-art performance on automatic speech recognition (ASR) and proved to be extremely useful in low label-resource settings. However, the success of SSL models has yet to transfer to utterance-level tas…

Cited by 7SourcePDFScholar
2021

A Comparative Study of Acoustic and Linguistic Features Classification for Alzheimer's Disease Detection

ICASSP 2021accepted

With the global population ageing rapidly, Alzheimer's disease (AD) is particularly prominent in older adults, which has an insidious onset followed by gradual, irreversible deterioration in cognitive domains (memory, communication, etc). Thus the detection of Alzheimer's disease is crucial for time…

Cited by 0SourceScholar
2020

Information Maximized Variational Domain Adversarial Learning for Speaker Verification

ICASSP 2020accepted

Domain mismatch is a common problem in speaker verification. This paper proposes an information-maximized variational domain adversarial neural network (InfoVDANN) to reduce domain mismatch by incorporating an InfoVAE into domain adversarial training (DAT). DAT aims to produce speaker discriminative…

Cited by 0SourceScholar
2020

Multi-Level Deep Neural Network Adaptation for Speaker Verification Using MMD and Consistency Regularization

ICASSP 2020accepted

Adapting speaker verification (SV) systems to a new environment is a very challenging task. Current adaptation methods in SV mainly focus on the backend, i.e, adaptation is carried out after the speaker embeddings have been created. In this paper, we present a DNN-based adaptation method using maxim…

Cited by 0SourceScholar
2019

Semi-supervised Nuisance-attribute Networks for Domain Adaptation

ICASSP 2019accepted

How to overcome the training and test data mismatch in speaker verification systems has been a focus of research recently. In this paper, we propose a semi-supervised nuisance attribute network (SNAN) to reduce the domain mismatch in i-vectors and x-vectors. SNANs are based on the idea of nuisance a…

Cited by 0SourceScholar