← Search

Li Fu

6 accepted papers

2025

UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition

ICASSP 2025accepted

Recent advancements in scaling up models have significantly improved performance in Automatic Speech Recognition (ASR) tasks. However, training large ASR models from scratch remains costly. To address this issue, we introduce UME, a novel method that efficiently Upcycles pretrained dense ASR checkpo…

Cited by 0SourceScholar
2024

Do Self-Supervised Speech and Language Models Extract Similar Representations as Human Brain?

ICASSP 2024accepted

Speech and language models trained through self-supervised learning (SSL) demonstrate strong alignment with brain activity during speech and language perception. However, given their distinct training modalities, it remains unclear whether they correlate with the same neural aspects. We directly add…

Cited by 0SourceScholar
2024

Neural2speech: A Transfer Learning Framework for Neural-Driven Speech Reconstruction

ICASSP 2024accepted

Reconstructing natural speech from neural activity is vital for enabling direct communication via brain-computer interfaces. Previous efforts have explored the conversion of neural recordings into speech using complex deep neural network (DNN) models trained on extensive neural recording data, which…

Cited by 0SourceScholar
2023

GPMO: Gradient Perturbation-Based Contrastive Learning for Molecule Optimization

IJCAI 2023poster

Optimizing molecules with desired properties is a crucial step in de novo drug design. While translation-based methods have achieved initial success, they continue to face the challenge of the “exposure bias” problem. The challenge of preventing the “exposure bias” problem of molecule optimization…

Cited by 5SourcePDFScholar
2023

UFO2: A Unified Pre-Training Framework for Online and Offline Speech Recognition

ICASSP 2023accepted

In this paper, we propose a Unified pre-training Framework for Online and Offline (UFO2) Automatic Speech Recognition (ASR), which 1) simplifies the two separate training workflows for online and offline modes into one process, and 2) improves the Word Error Rate (WER) performance with limited utter…

Cited by 0SourceScholar