← Search

Alexander Waibel

17 accepted papers

2026

A COCKTAIL-PARTY BENCHMARK: MULTI-MODAL DATASET AND COMPARATIVE EVALUATION RESULTS

ICASSP 2026poster

We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a single-room setting using audio, visual, and contextual cues. MCoRec captures natural multi-party conversations where the…

Cited by 0SourcePDFScholar
2026

A MULTIMODAL DEPTH-AWARE METHOD FOR EMBODIED REFERENCE UNDERSTANDING

ICASSP 2026poster

Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While prior works have shown progress in open-vocabulary object detection, they often fail in ambiguous scenarios where multiple candidate objects exist in t…

Cited by 0SourcePDFScholar
2026

ASSESSING IDENTITY LEAKAGE IN TALKING FACE GENERATION: METRICS AND EVALUATION FRAMEWORK

ICASSP 2026poster

Video editing-based talking face generation aims to preserve video details such as pose, lighting, and gestures while modifying only lip motion, often using an identity reference image to maintain speaker consistency. However, this mechanism can introduce lip leakage, where generated lips are influe…

Cited by 0SourcePDFScholar
2025

Factorized-VITS: Decoupling Prosody and Text in End-to-End Speech Synthesis without External or Secondary Aligner

ICASSP 2025accepted

We propose Factorized-VITS, an advanced end-to-end text-to-speech model that incorporates explicit text-side prosody modeling control into VITS while achieving a clean factorization of the audio prior hidden space into text and prosody subspaces. Unlike previous works that rely on external or second…

Cited by 0SourceScholar
2025

Improving Pronunciation and Accent Conversion through Knowledge Distillation And Synthetic Ground-Truth from Native TTS

ICASSP 2025accepted

Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciation issues, which can make it difficult for listeners to understand them. Hence,…

Cited by 0SourceScholar
2024

Audio-driven Talking Face Generation with Stabilized Synchronization Loss

ECCV 2024poster

"Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper, we start by identifying several issues with existing synchronization learning…

Cited by 1SourcePDFScholar
2024

DECM: Evaluating Bilingual ASR Performance on a Code-switching/mixing Benchmark

COLING 2024main

Automatic Speech Recognition has made significant progress, but challenges persist. Code-switched (CSW) Speech presents one such challenge, involving the mixing of multiple languages by a speaker. Even when multilingual ASR models are trained, each utterance on its own usually remains monolingual. W…

2024

Decoupled Vocabulary Learning Enables Zero-Shot Translation from Unseen Languages

ACL 2024long

Multilingual neural machine translation systems learn to map sentences of different languages into a common representation space. Intuitively, with a growing number of seen languages the encoder sentence representation grows more flexible and easily adaptable to new languages. In this work, we test…

2024

SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading

EMNLP 2024main

With the rapid development of Large Language Models (LLMs), it is crucial to have benchmarks which can evaluate the ability of LLMs on different domains. One common use of LLMs is performing tasks on scientific topics, such as writing algorithms, querying databases or giving mathematical proofs. Ins…

2023

AdapITN: A Fast, Reliable, and Dynamic Adaptive Inverse Text Normalization

ICASSP 2023accepted

Inverse text normalization (ITN) is the task that transforms text in spoken-form into written-form. While automatic speech recognition (ASR) produces text in spoken-form, human and natural language understanding systems prefer to consume text in written-form. ITN generally deals with semiotic phrase…

Cited by 0SourceScholar
2023

SYNTACC : Synthesizing Multi-Accent Speech By Weight Factorization

ICASSP 2023accepted

Conventional multi-speaker text-to-speech synthesis (TTS) is known to be capable of synthesizing speech for multiple voices, yet it cannot generate speech in different accents. This limitation has motivated us to develop SYNTACC (Synthesizing speech with accents) which adapts conventional multi-spea…

Cited by 0SourceScholar
2016

An empirical exploration of CTC acoustic models

ICASSP 2016accepted

The connectionist temporal classification (CTC) loss function has several interesting properties relevant for automatic speech recognition (ASR): applied on top of deep recurrent neural networks (RNNs), CTC learns the alignments between speech frames and label sequences automatically, which removes…

Cited by 0SourceScholar