← Search

Jia Qi Yip

11 accepted papers

2026

SumRA: Parameter Efficient Fine-tuning with Singular Value Decomposition and Summed Orthogonal Basis

ICLR 2026poster

Parameter-efficient fine-tuning (PEFT) aims to adapt large pretrained speech models using fewer trainable parameters while maintaining performance. Low-Rank Adaptation (LoRA) achieves this by decomposing weight updates into two low-rank matrices, $A$ and $B$, such that $W'=W_0+BA$. Previous studies…

Cited by 0SourceScholar
2025

Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

ICLR 2025poster

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication…

2025

Extending Whisper for Emotion Prediction Using Word-level Pseudo Labels

ICASSP 2025accepted

This paper extends Whisper’s automatic speech recognition (ASR) capabilities to perform speech-based emotion recognition (SER) by incorporating word-level emotion classification alongside ASR output. We generate four emotion pseudo-labels (neutral, happy, sad, angry) for each word using a pretrained…

Cited by 0SourceScholar
2025

Speech Enhancement Using Continuous Embeddings of Neural Audio Codec

ICASSP 2025accepted

Recent advancements in Neural Audio Codec (NAC) models have inspired their use in various speech processing tasks, including speech enhancement (SE). In this work, we propose a novel, efficient SE approach by leveraging the pre-quantization output of a pretrained NAC encoder. Unlike prior NAC-based…

Cited by 0SourceScholar
2025

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

NAACL 2025system demonstrations

In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface with flexible configuration and dependency control, making it user-friendly and efficient. With full installation, VERSA of…

2024

Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification

ICASSP 2024accepted

Knowledge distillation (KD) is used to enhance automatic speaker verification performance by ensuring consistency between large teacher networks and lightweight student networks at the embedding level or label level. However, the conventional label-level KD overlooks the significant knowledge from n…

Cited by 0SourceScholar
2024

MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation

ICASSP 2024accepted

Our previously proposed MossFormer has achieved promising performance in monaural speech separation. However, it predominantly adopts a self-attention-based MossFormer module, which tends to emphasize longer-range, coarser-scale dependencies, with a deficiency in effectively modelling finer-scale re…

Cited by 0SourceScholar
2024

SPGM: Prioritizing Local Features for Enhanced Speech Separation Performance

ICASSP 2024accepted

Dual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, w…

Cited by 0SourceScholar
2023

Contrastive Speech Mixup for Low-Resource Keyword Spotting

ICASSP 2023accepted

Most of the existing neural-based models for keyword spotting (KWS) in smart devices require thousands of training samples to learn a decent audio representation. However, with the rising demand for smart devices to become more person-alized, KWS models need to adapt quickly to smaller user samples.…

Cited by 15SourceScholar
2023

De'hubert: Disentangling Noise in a Self-Supervised Model for Robust Speech Recognition

ICASSP 2023accepted

Existing self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However, many current models are trained on a clean corpus from a single source, which tends to do poorly when noise is present d…

Cited by 0SourceScholar