← Search

Alexandros Haliassos

8 accepted papers

2026

Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition

ICLR 2026poster

Unified Speech Recognition (USR) has emerged as a semi-supervised framework for training a single model for audio, visual, and audiovisual speech recognition, achieving state-of-the-art results on in-distribution benchmarks. However, its reliance on autoregressive pseudo-labelling makes training exp…

Cited by 0SourcecodeScholar
2024

BRAVEn: Improving Self-supervised pre-training for Visual and Auditory Speech Recognition

ICASSP 2024accepted

Self-supervision has recently shown great promise for learning visual and auditory speech representations from unlabelled data. In this work, we propose BRAVEn, an extension to the recent RAVEn method, which learns speech representations entirely from raw audio-visual data. Our modifications to RAVE…

Cited by 0SourceScholar
2024

Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs

NeurIPS 2024poster

Research in auditory, visual, and audiovisual speech recognition (ASR, VSR, and AVSR, respectively) has traditionally been conducted independently. Even recent self-supervised studies addressing two or all three tasks simultaneously tend to yield separate models, leading to disjoint inference pipeli…

2023

Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels

ICASSP 2023accepted

Audio-visual speech recognition has received a lot of attention due to its robustness against acoustic noise. Recently, the performance of automatic, visual, and audio-visual speech recognition (ASR, VSR, and AV-ASR, respectively) has been substantially improved, mainly due to the use of larger mode…

Cited by 0SourceScholar
2023

Jointly Learning Visual and Auditory Speech Representations from Raw Data

ICLR 2023poster

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets generated by slowly-evolving momentum encoders. Driven by the inherent differen…

2023

Learning Cross-Lingual Visual Speech Representations

ICASSP 2023accepted

Cross-lingual self-supervised learning has been a growing research topic in the last few years. However, current works only explored the use of audio signals to create representations. In this work, we study cross-lingual self-supervised visual representation learning. We use the recently-proposed R…

Cited by 0SourceScholar
2022

Leveraging Real Talking Faces via Self-Supervision for Robust Forgery Detection

CVPR 2022poster

One of the most pressing challenges for the detection of face-manipulated videos is generalising to forgery methods not seen during training while remaining effective under common corruptions such as compression. In this paper, we examine whether we can tackle this issue by harnessing videos of real…

Cited by 149PDFcodeScholar
2021

Lips Don't Lie: A Generalisable and Robust Approach To Face Forgery Detection

CVPR 2021poster

Although current deep learning-based face forgery detectors achieve impressive performance in constrained scenarios, they are vulnerable to samples created by unseen manipulation methods. Some recent works show improvements in generalisation but rely on cues that are easily corrupted by common post-…

Cited by 515PDFcodeScholar