← Search

Tobias Bocklet

12 accepted papers

2025

Adapter-Based Multi-Agent AVSR Extension for Pre-Trained ASR Models

ICASSP 2025accepted

We present an approach to Audio-Visual Speech Recognition that builds on a pre-trained Whisper model. To infuse visual information into this audio-only model, we extend it with an AV fusion module and LoRa adapters, one of the most up-to-date adapter approaches. One advantage of adapter-based approa…

Cited by 0SourceScholar
2025

Digital Operating Mode Classification of Real-World Amateur Radio Transmissions

ICASSP 2025accepted

This study presents an ML approach for classifying digital radio operating modes evaluated on real-world transmissions. We generated 98 different parameterized radio signals from 17 digital operating modes, transmitted each of them on the 70 cm (UHF) amateur radio band, and recorded our transmission…

Cited by 0SourceScholar
2025

FedSVD: Adaptive Orthogonalization for Private Federated Learning with LoRA

NeurIPS 2025poster

Low-Rank Adaptation (LoRA), which introduces a product of two trainable low-rank matrices into frozen pre-trained weights, is widely used for efficient fine-tuning of language models in federated learning (FL). However, when combined with differentially private stochastic gradient descent (DP-SGD),…

Cited by 0SourceScholar
2025

Optimized Self-supervised Training with BEST-RQ for Speech Recognition

ICASSP 2025accepted

Self-supervised learning has been successfully used for various speech related tasks, including automatic speech recognition. BERT-based Speech pre-Training with Random-projection Quantizer (BEST-RQ) has achieved state-of-the-art results in speech recognition. In this work, we further optimize the B…

Cited by 3SourceScholar
2025

SafeRoute: Adaptive Model Selection for Efficient and Accurate Safety Guardrails in Large Language Models

ACL 2025finding

Deploying large language models (LLMs) in real-world applications requires robust safety guard models to detect and block harmful user prompts. While large safety guard models achieve strong performance, their computational cost is substantial. To mitigate this, smaller distilled models are used, bu…

2024

Optimized Speculative Sampling for GPU Hardware Accelerators

EMNLP 2024main

In this work, we optimize speculative sampling for parallel hardware accelerators to improve sampling speed. We notice that substantial portions of the intermediate matrices necessary for speculative sampling can be computed concurrently. This allows us to distribute the workload across multiple GPU…

2024

Towards Interpretability of Automatic Phoneme Analysis in Cleft Lip and Palate Speech

ICASSP 2024accepted

Cleft Lip and Palate ranks among the most common congenital abnormalities and significantly influences speech articulation, resulting in varying phonemic impacts. In a clinical context, a detailed diagnosis is carried out by time-consuming perceptual evaluations. We use perceptual ratings of differe…

Cited by 0SourceScholar
2021

Acoustic and Linguistic Analyses to Assess Early-Onset and Genetic Alzheimer's Disease

ICASSP 2021accepted

The PSEN1-E280A or Paisa mutation is responsible for most of Early-Onset Alzheimer’s (EOA) disease cases in Colombia. It affects a large kindred of over 5000 members that present the same phenotype. The most common symptoms are related to language disorders, where speech fluency is also affected due…

Cited by 0SourceScholar
2020

Comparison of User Models Based on GMM-UBM and I-Vectors for Speech, Handwriting, and Gait Assessment of Parkinson's Disease Patients

ICASSP 2020accepted

Parkinson's disease is a neurodegenerative disorder characterized by the presence of different motor impairments. Information from speech, handwriting, and gait signals have been considered to evaluate the neurological state of the patients. On the other hand, user models based on Gaussian mixture m…

Cited by 0SourceScholar
2017

Multi-view representation learning via gcca for multimodal analysis of Parkinson's disease

ICASSP 2017accepted

Information from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized…

Cited by 35SourceScholar
2017

On the impact of non-modal phonation on phonological features

ICASSP 2017accepted

Different modes of vibration of the vocal folds contribute significantly to the voice quality. The neutral mode phonation, often used in a modal voice, is one against which the other modes can be contrastively described, also called non-modal phonations. This paper investigates the impact of non-mod…

Cited by 0SourceScholar