← Search

Tanja Schultz

18 accepted papers

2025

ATGnet: Adaptive Temporal Graph Network for EEG-enabled Sound Source Tracking in Cocktail Party Scenarios

ICASSP 2025accepted

Decoding selective auditory attention from electroencephalography (EEG) signals has gained considerable interest. However, few studies have looked into tracking the dynamic trajectory of moving sound source in complex auditory environments, e.g. with multiple moving speakers. We propose a novel mode…

Cited by 0SourceScholar
2025

Shabdh: A multi lingual zero-shot voice cloning approach with speaker disentanglement

ICASSP 2025accepted

This paper presents a zero-shot voice cloning system leveraging the DIS-Vector framework, which disentangles and encodes key speech features: content, pitch, timbre, and rhythm. Using the YourTTS architecture, the system synthesizes high-quality speech with precise control over both speaker identity…

Cited by 0SourceScholar
2025

Speech Separation for Low-Resource Languages

ICASSP 2025accepted

Speech separation aims to equip machines with the human ability of selective listening, i.e. to focus attention on specific information in spoken communication. Studies have shown that the language spoken in a cocktail party scenario matters. While the development of speech separation models can lev…

Cited by 0SourceScholar
2023

Multi-Head Attention and GRU for Improved Match-Mismatch Classification of Speech Stimulus and EEG Response

ICASSP 2023accepted

This work is based on the participation by the HyperAttention team in the Auditory EEG Decoding Challenge, 2023 (ICASSP 2023 Signal Processing Grand Challenge) task 1, which deals with the match-mismatch classification of speech stimuli and EEG responses of human listeners. We demonstrate the benefi…

Cited by 0SourceScholar
2023

Multi-Speaker Speech Synthesis from Electromyographic Signals by Soft Speech Unit Prediction

ICASSP 2023accepted

Electromyographic (EMG) signals of articulatory muscles reflect the speech production process even if the user is speaking silently i.e. moving the articulators without producing audible sound. We propose Speech-Unit-based EMG-to-Speech (SU-E2S), a system which relies on EMG to synthesize speech whi…

Cited by 0SourceScholar
2022

An Overview of the FIRST ICASSP Special Session on Computer Audition for Healthcare

ICASSP 2022accepted

Audio has been increasingly used as a novel digital phenotype that carries important information of the subject’s health status. We can find tremendous efforts given to this young and promising field, i.e., computer audition for healthcare (CA4H), whereas the application scenarios have not been full…

Cited by 0SourceScholar
2022

Experts Versus All-Rounders: Target Language Extraction for Multiple Target Languages

ICASSP 2022accepted

Target language extraction (TLE) is a novel task in the field of selective auditory attention, which seeks to extract all speech signals that are spoken in a target language from other sources in a multilingual cocktail party. In our prior studies, a TLE model was trained to extract a predefined, si…

Cited by 0SourceScholar
2022

Exploring Dementia Detection from Speech: Cross Corpus Analysis

ICASSP 2022accepted

In this work, we present a qualitative and quantitative analysis of speech and language features derived from two different corpora with the aim to predict early signs of dementia. One corpus consists of the Interdisciplinary Longitudinal Study on Adult Development and Aging (ILSE) designed to inves…

Cited by 22SourceScholar
2022

Hybrid sub-word segmentation for handling long tail in morphologically rich low resource languages

ICASSP 2022accepted

Dealing with Out Of Vocabulary (OOV) words or unseen words is one of the main issues of Machine Translation (MT) as well as automatic speech recognition (ASR) systems. For morphologically rich languages having high type token ratio, the OOV percentage is also quite high. Sub-word segmentation has be…

Cited by 0SourceScholar
2022

Towards Closed-Loop Speech Synthesis from Stereotactic EEG: A Unit Selection Approach

ICASSP 2022accepted

Neurological disorders can severely impact speech communication. Recently, neural speech prostheses have been proposed that reconstruct intelligible speech from neural signals recorded superficially on the cortex. Thus far, it has been unclear whether similar reconstruction is feasible from deeper b…

Cited by 0SourceScholar
2021

End-to-End Multilingual Automatic Speech Recognition for Less-Resourced Languages: The Case of Four Ethiopian Languages

ICASSP 2021accepted

The End-to-End (E2E) approach, which maps a sequence of input features into a sequence of graphemes or words, to Automatic Speech Recognition (ASR) is a hot research agenda. It is interesting for less-resourced languages since it avoids the use of pronunciation dictionary, which is one of the major…

Cited by 0SourceScholar
2020

DNN-Based Speech Recognition for Globalphone Languages

ICASSP 2020accepted

This paper describes new reference benchmark results based on hybrid Hidden Markov Model and Deep Neural Networks (HMM-DNN) for the GlobalPhone (GP) multilingual text and speech database. GP is a multilingual database of high-quality read speech with corresponding transcriptions and pronunciation di…

Cited by 0SourceScholar
2020

Deep Neural Networks Based Automatic Speech Recognition for Four Ethiopian Languages

ICASSP 2020accepted

In this work, we present speech recognition systems for four Ethiopian languages: Amharic, Tigrigna, Oromo and Wolaytta. We have used comparable training corpora of about 20 to 29 hours speech and evaluation speech of about 1 hour for each of the languages. For Amharic and Tigrigna, lexical and lang…

Cited by 0SourceScholar
2020

From Human to Robot Everyday Activity

IROS 2020poster

The Everyday Activities Science and Engineering (EASE) Collaborative Research Consortium’s mission to enhance the performance of cognition-enabled robots establishes its foundation in the EASE Human Activities Data Analysis Pipeline. Through collection of diverse human activity information resources…

Cited by 11SourceScholar
2015

Cross-lingual lexical language discovery from audio data using multiple translations

ICASSP 2015accepted

Zero-resource Automatic Speech Recognition (ZR ASR) addresses target languages without given pronunciation dictionary, transcribed speech, and language model. Lexical discovery for ZR ASR aims to extract word-like chunks from speech. Lexical discovery benefits from the availability of written transl…

Cited by 0SourceScholar
2015

Telemanipulation with force-based display of proximity fields

IROS 2015poster

In this paper we show and evaluate the design of a novel telemanipulation system that maps proximity values, acquired inside of a gripper, to forces a user can feel through a haptic input device. The command console is complemented by input-devices that give the user an intuitive control over parame…

Cited by 14SourceScholar