← Search

Stavros Petridis

34 accepted papers

2026

FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs

CVPR 2026

We present FlashLips, a two-stage, mask-free lip-sync system that decouples lips control from rendering and achieves real-time performance, with our U-Net variant running at over 100 FPS on a single GPU, while matching the visual quality of larger state-of-the-art models. Stage 1 is a compact, one-s

Cited by 0SourceScholar
2026

OMNI-AVSR: TOWARDS UNIFIED MULTIMODAL SPEECH RECOGNITION WITH LARGE LANGUAGE MODELS

ICASSP 2026poster

Large language models (LLMs) have recently achieved impressive results in speech recognition across multiple modalities, including Auditory Speech Recognition (ASR), Visual Speech Recognition (VSR), and Audio-Visual Speech Recognition (AVSR). Despite this progress, current LLM-based approaches typic…

Cited by 0SourcePDFScholar
2026

Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech Recognition

ICLR 2026poster

Unified Speech Recognition (USR) has emerged as a semi-supervised framework for training a single model for audio, visual, and audiovisual speech recognition, achieving state-of-the-art results on in-distribution benchmarks. However, its reliance on autoregressive pseudo-labelling makes training exp…

Cited by 0SourcecodeScholar
2025

Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction

ICASSP 2025accepted

In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Contextual Speech Extraction (CSE). Unlike traditional TSE methods that rely on pre-recorded enrollment utterances, video o…

Cited by 0SourceScholar
2025

Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models

ICASSP 2025accepted

This paper investigates the under-explored area of low-rank weight training for large-scale Conformer-based speech recognition models from scratch. Our study demonstrates the viability of this training paradigm for such models, yielding several notable findings. Firstly, we discover that applying a…

Cited by 0SourceScholar
2025

KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation

CVPR 2025poster

Current audio-driven facial animation methods achieve impressive results for short videos but suffer from error accumulation and identity drift when extended to longer durations. Existing methods attempt to mitigate this through external spatial control, increasing long-term consistency but compromi…

Cited by 2SourcePDFScholar
2025

Large Language Models are Strong Audio-Visual Speech Recognition Learners

ICASSP 2025accepted

Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the au…

Cited by 0SourceScholar
2025

MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition

NeurIPS 2025poster

Large language models (LLMs) have recently shown strong potential in audio-visual speech recognition (AVSR), but their high computational demands and sensitivity to token granularity limit their practicality in resource-constrained settings. Token compression methods can reduce inference cost, but t…

Cited by 0SourceScholar
2025

Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

ICCV 2025poster

We explore a novel zero-shot Audio-Visual Speech Recognition (AVSR) framework, dubbed Zero-AVSR, which enables speech recognition in target languages without requiring any audio-visual speech data in those languages. Specifically, we introduce the Audio-Visual Speech Romanizer (AV-Romanizer), which…

2024

BRAVEn: Improving Self-supervised pre-training for Visual and Auditory Speech Recognition

ICASSP 2024accepted

Self-supervision has recently shown great promise for learning visual and auditory speech representations from unlabelled data. In this work, we propose BRAVEn, an extension to the recent RAVEn method, which learns speech representations entirely from raw audio-visual data. Our modifications to RAVE…

Cited by 0SourceScholar
2024

EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars

CVPR 2024poster

Head avatars animated by visual signals have gained popularity particularly in cross-driving synthesis where the driver differs from the animated character a challenging but highly practical approach. The recently presented MegaPortraits model has demonstrated state-of-the-art results in this domain…

Cited by 26SourcePDFScholar
2024

Hearing Loss Detection From Facial Expressions in One-On-One Conversations

ICASSP 2024accepted

Individuals with impaired hearing experience difficulty in conversations, especially in noisy environments. This difficulty often manifests as a change in behavior and may be captured via facial expressions, such as the expression of discomfort or fatigue. In this work, we build on this idea and int…

Cited by 0SourceScholar
2024

Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs

NeurIPS 2024poster

Research in auditory, visual, and audiovisual speech recognition (ASR, VSR, and AVSR, respectively) has traditionally been conducted independently. Even recent self-supervised studies addressing two or all three tasks simultaneously tend to yield separate models, leading to disjoint inference pipeli…

2023

Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels

ICASSP 2023accepted

Audio-visual speech recognition has received a lot of attention due to its robustness against acoustic noise. Recently, the performance of automatic, visual, and audio-visual speech recognition (ASR, VSR, and AV-ASR, respectively) has been substantially improved, mainly due to the use of larger mode…

Cited by 0SourceScholar
2023

Jointly Learning Visual and Auditory Speech Representations from Raw Data

ICLR 2023poster

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets generated by slowly-evolving momentum encoders. Driven by the inherent differen…

2023

LA-VOCE: LOW-SNR Audio-Visual Speech Enhancement Using Neural Vocoders

ICASSP 2023accepted

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker’s lip movements. This approach has been shown to yield improvements over audio-only speech enhancement, particularly for the removal of interferin…

Cited by 0SourceScholar
2023

Learning Cross-Lingual Visual Speech Representations

ICASSP 2023accepted

Cross-lingual self-supervised learning has been a growing research topic in the last few years. However, current works only explored the use of audio signals to create representations. In this work, we study cross-lingual self-supervised visual representation learning. We use the recently-proposed R…

Cited by 0SourceScholar
2023

SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision

CVPR 2023poster

Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video datasets are limited in size. In this paper, for the first time, we study the potential of leveraging synthetic visual…

Cited by 27SourcePDFScholar
2022

Leveraging Real Talking Faces via Self-Supervision for Robust Forgery Detection

CVPR 2022poster

One of the most pressing challenges for the detection of face-manipulated videos is generalising to forgery methods not seen during training while remaining effective under common corruptions such as compression. In this paper, we examine whether we can tackle this issue by harnessing videos of real…

Cited by 149PDFcodeScholar
2021

DINO: A Conditional Energy-Based GAN for Domain Translation

ICLR 2021poster

Domain translation is the process of transforming data from one domain to another while preserving the common semantics. Some of the most popular domain translation systems are based on conditional generative adversarial networks, which use source domain data to drive the generator and as an input t…

2021

Lips Don't Lie: A Generalisable and Robust Approach To Face Forgery Detection

CVPR 2021poster

Although current deep learning-based face forgery detectors achieve impressive performance in constrained scenarios, they are vulnerable to samples created by unseen manipulation methods. Some recent works show improvements in generalisation but rely on cues that are easily corrupted by common post-…

Cited by 515PDFcodeScholar
2021

Towards Practical Lipreading with Distilled and Efficient Models

ICASSP 2021accepted

Lipreading has witnessed a lot of progress due to the resurgence of neural networks. Recent works have placed emphasis on aspects such as improving performance by finding the optimal architecture or improving generalization. However, there is still a significant gap between the current methodologies…

Cited by 0SourceScholar
2020

Speech-Driven Facial Animation Using Polynomial Fusion of Features

ICASSP 2020accepted

Speech-driven facial animation involves using a speech signal to generate realistic videos of talking faces. Recent deep learning approaches to facial synthesis rely on extracting low-dimensional representations and concatenating them, followed by a decoding step of the concatenated vector. This acc…

Cited by 0SourceScholar
2020

Towards Pose-Invariant Lip-Reading

ICASSP 2020accepted

Lip-reading models have been significantly improved recently thanks to powerful deep learning architectures. However, most works focused on frontal or near frontal views of the mouth. As a consequence, lip-reading performance seriously deteriorates in non-frontal mouth views. In this work, we presen…

Cited by 0SourceScholar
2020

Visually Guided Self Supervised Learning of Speech Representations

ICASSP 2020accepted

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very limited work that studies the interaction between the two modaliti…

Cited by 0SourceScholar
2018

End-to-End Audiovisual Speech Recognition

ICASSP 2018accepted

Several end-to-end deep learning approaches have been recently presented which extract either audio or visual features from the input images or audio signals and perform speech recognition. However, research on end-to-end audiovisual models is very limited. In this work, we present an end-to-end aud…

Cited by 0SourceScholar
2017

Audio-visual object localization and separation using low-rank and sparsity

ICASSP 2017accepted

The ability to localize visual objects that are associated with an audio source and at the same time seperate the audio signal is a corner stone in several audio-visual signal processing applications. Past efforts usually focused on localizing only the visual objects, without audio separation abilit…

Cited by 0SourceScholar