← Search

Dorothea Kolossa

15 accepted papers

2024

DistriBlock: Identifying adversarial audio samples by leveraging characteristics of the output distribution

UAI 2024poster

Adversarial attacks can mislead automatic speech recognition (ASR) systems into predicting an arbitrary target text, thus posing a clear security threat. To prevent such attacks, we propose DistriBlock, an efficient detection strategy applicable to any ASR system that predicts a probability distribu…

2024

Who Wrote When? Author Diarization in Social Media Discussions

EMNLP 2024finding

We are proposing a novel framework for author diarization, i.e. attributing comments in online discussions to individual authors. We consider an innovative approach that merges pre-trained neural representations of writing style with author-conditional encoder-decoder diarization, enhanced by a Cond…

2021

Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial Domain

ICASSP 2021accepted

Estimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamforming or assigning unique speaker identities. Recently, several approaches utili…

Cited by 0SourceScholar
2021

Fusing Information Streams in End-to-End Audio-Visual Speech Recognition

ICASSP 2021accepted

End-to-end acoustic speech recognition has quickly gained widespread popularity and shows promising results in many studies. Specifically the joint transformer/CTC model pro-vides very good performance in many tasks. However, under noisy and distorted conditions, the performance still degrades notab…

Cited by 0SourceScholar
2020

A Dynamic Stream Weight Backprop Kalman Filter for Audiovisual Speaker Tracking

ICASSP 2020accepted

Audiovisual speaker tracking is an application that has been tackled by a wide range of classical approaches based on Gaussian filters, most notably the well-known Kalman filter. Recently, a specific Kalman filter implementation was proposed for this task, which incorporated dynamic stream weights t…

Cited by 0SourceScholar
2020

Leveraging Frequency Analysis for Deep Fake Image Recognition

ICML 2020poster

Deep neural networks can generate images that are astonishingly realistic, so much so that it is often hard for humans to distinguish them from actual photos. These achievements have been largely made possible by Generative Adversarial Networks (GANs). While deep fake images have been thoroughly inv…

2020

Variational Autoencoder with Embedded Student-t Mixture Model for Authorship Attribution

COLING 2020main

Traditional computational authorship attribution describes a classification task in a closed-set scenario. Given a finite set of candidate authors and corresponding labeled texts, the objective is to determine which of the authors has written another set of anonymous or disputed texts. In this work,…

Cited by 2SourcePDFScholar
2019

Learning Dynamic Stream Weights for Linear Dynamical Systems Using Natural Evolution Strategies

ICASSP 2019accepted

Multimodal data fusion is an important aspect of many object localization and tracking frameworks that rely on sensory observations from different sources. A prominent example is audiovisual speaker localization, where the incorporation of visual information has shown to benefit overall performance,…

Cited by 0SourceScholar
2019

Similarity Learning for Authorship Verification in Social Media

ICASSP 2019accepted

Authorship verification tries to answer the question if two documents with unknown authors were written by the same author or not. A range of successful technical approaches has been proposed for this task, many of which are based on traditional linguistic features such as n-grams. These algorithms…

Cited by 0SourceScholar
2018

Potential-Field-Based Active Exploration for Acoustic Simultaneous Localization and Mapping

ICASSP 2018accepted

This paper presents a novel framework for active exploration in the context of acoustic simultaneous localization and mapping (SLAM) using a microphone array mounted on a mobile robotic agent. Acoustic SLAM aims at building a map of acoustic sources present in the environment and simultaneously esti…

Cited by 0SourceScholar
2017

Improving audio-visual speech recognition using deep neural networks with dynamic stream reliability estimates

ICASSP 2017accepted

Audio-visual speech recognition is a promising approach to tackling the problem of reduced recognition rates under adverse acoustic conditions. However, finding an optimal mechanism for combining multi-modal information remains a challenging task. Various methods are applicable for integrating acous…

Cited by 0SourceScholar
2017

Monte Carlo exploration for active binaural localization

ICASSP 2017accepted

This study introduces a machine hearing system for robot audition, which enables a robotic agent to pro-actively minimize the uncertainty of sound source location estimates through motion. The proposed system is based on an active exploration approach, providing a means to model and predict effects…

Cited by 0SourceScholar
2017

Speaker localization in reverberant rooms based on direct path dominance test statistics

ICASSP 2017accepted

Speaker localization using microphone arrays is typically based on the expected phase and amplitude differences between microphones as a function of the wave arrival direction. However, in rooms with significant reverberation, the direct sound is contaminated by reflections and localization often fa…

Cited by 35SourceScholar
2016

Robust audiovisual speech recognition using noise-adaptive linear discriminant analysis

ICASSP 2016accepted

Automatic speech recognition (ASR) has become a widespread and convenient mode of human-machine interaction, but it is still not sufficiently reliable when used under highly noisy or reverberant conditions. One option for achieving far greater robustness is to include another modality that is unaffe…

Cited by 0SourceScholar
2016

Twin-HMM-based non-intrusive speech intelligibility prediction

ICASSP 2016accepted

Most of the objective measures employed for speech intelligibility prediction require a clean reference signal, which is not accessible in all realistic scenarios. In this paper, we propose to re-synthesize the relevant features of the clean signal using only the noisy speech signal and utilize them…

Cited by 0SourceScholar