← Search

Philip N. Garner

10 accepted papers

2025

Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity

ICLR 2025poster

Architectures such as Linformer and Mamba have recently emerged as competitive linear time replacements for transformers. However, corresponding large pretrained models are often unavailable, especially in non-text domains. To remedy this, we present a Cross-Architecture Layerwise Distillation (CALD…

2023

The Interpreter Understands Your Meaning: End-to-end Spoken Language Understanding Aided by Speech Translation

EMNLP 2023long findings

End-to-end spoken language understanding (SLU) remains elusive even with current large pretrained language models on text and speech, especially in multilingual cases. Machine translation has been established as a powerful pretraining objective on text as it enables the model to capture high-level s…

Cited by 0SourcecodeScholar
2019

An End-to-end Network to Synthesize Intonation Using a Generalized Command Response Model

ICASSP 2019accepted

The generalized command response (GCR) model represents intonation as a superposition of muscle responses to spike command signals. We have previously shown that the spikes can be predicted by a two-stage system, consisting of a recurrent neural network and a post-processing procedure, but the respo…

Cited by 0SourceScholar
2019

Empirical Evaluation and Combination of Punctuation Prediction Models Applied to Broadcast News

ICASSP 2019accepted

Natural language processing techniques are dependent upon punctuation to work well. When their input is taken from speech recognition, it is necessary to reconstruct the punctuation; in particular sentence boundaries. We define a range of features from low level acoustics to those with high level le…

Cited by 0SourceScholar
2015

Robust microphone placement for source localization from noisy distance measurements

ICASSP 2015accepted

We propose a novel algorithm to design an optimum array geometry for source localization inside an enclosure. We assume a square-law decay propagation model for the sound acquisition so that the additive noise on the measured source-microphone distances is proportional to the distances regardless of…

Cited by 0SourceScholar