← Search

Jing Pan

6 accepted papers

2025

CogniDual Framework: Self-Training Large Language Models within a Dual-System Theoretical Framework for Improving Cognitive Tasks

ICASSP 2025accepted

Cognitive psychology investigates perception, attention, memory, language, problem-solving, decision-making, and reasoning. Kahneman’s dual-system theory elucidates the human decision-making process, distinguishing between the rapid, intuitive System 1 and the deliberative, rational System 2. Recent…

Cited by 0SourceScholar
2025

Target word activity detector: An approach to obtain ASR word boundaries without lexicon

ICASSP 2025accepted

Obtaining word timestamp information from end-to-end (E2E) ASR models remains challenging due to the lack of explicit time alignment during training. This issue is further complicated in multilingual models. Existing methods, either rely on lexicons or introduce additional tokens, leading to scalabi…

Cited by 0SourceScholar
2024

WavLLM: Towards Robust and Adaptive Speech Large Language Model

EMNLP 2024finding

Recent advancements in large language models (LLMs) have expanded their scope in natural language processing (NLP) to encompass multimodal functions. However, integrating listening capabilities effectively remains a significant challenge for generalization and complex auditory task execution. In thi…

2022

Performance-Efficiency Trade-Offs in Unsupervised Pre-Training for Speech Recognition

ICASSP 2022accepted

This paper is a study of performance-efficiency trade-offs in pre-trained models for automatic speech recognition (ASR). We focus on wav2vec 2.0, and formalize several architecture designs that influence both the model performance and its efficiency. Putting together all our observations, we introdu…

Cited by 0SourceScholar
2022

SRU++: Pioneering Fast Recurrence with Attention for Speech Recognition

ICASSP 2022accepted

The Transformer architecture has been well adopted as a dominant architecture in most sequence transduction tasks including automatic speech recognition (ASR), since its attention mechanism excels in capturing long-range dependencies. While models built solely upon attention can be better paralleliz…

Cited by 0SourceScholar
2021

Multistream CNN for Robust Acoustic Modeling

ICASSP 2021accepted

This paper proposes multistream CNN, a novel neural network architecture for robust acoustic modeling in speech recognition tasks. The proposed architecture processes input speech with diverse temporal resolutions by applying different dilation rates to convolutional neural networks across multiple…

Cited by 0SourceScholar