← Search

Maja Pantic

46 accepted papers

2026

FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANs

CVPR 2026

We present FlashLips, a two-stage, mask-free lip-sync system that decouples lips control from rendering and achieves real-time performance, with our U-Net variant running at over 100 FPS on a single GPU, while matching the visual quality of larger state-of-the-art models. Stage 1 is a compact, one-s

Cited by 0SourceScholar
2025

Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction

ICASSP 2025accepted

In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Contextual Speech Extraction (CSE). Unlike traditional TSE methods that rely on pre-recorded enrollment utterances, video o…

Cited by 0SourceScholar
2025

Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models

ICASSP 2025accepted

This paper investigates the under-explored area of low-rank weight training for large-scale Conformer-based speech recognition models from scratch. Our study demonstrates the viability of this training paradigm for such models, yielding several notable findings. Firstly, we discover that applying a…

Cited by 0SourceScholar
2025

KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation

CVPR 2025poster

Current audio-driven facial animation methods achieve impressive results for short videos but suffer from error accumulation and identity drift when extended to longer durations. Existing methods attempt to mitigate this through external spatial control, increasing long-term consistency but compromi…

Cited by 2SourcePDFScholar
2025

Large Language Models are Strong Audio-Visual Speech Recognition Learners

ICASSP 2025accepted

Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the au…

Cited by 0SourceScholar
2025

MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition

NeurIPS 2025poster

Large language models (LLMs) have recently shown strong potential in audio-visual speech recognition (AVSR), but their high computational demands and sensitivity to token granularity limit their practicality in resource-constrained settings. Token compression methods can reduce inference cost, but t…

Cited by 0SourceScholar
2024

BRAVEn: Improving Self-supervised pre-training for Visual and Auditory Speech Recognition

ICASSP 2024accepted

Self-supervision has recently shown great promise for learning visual and auditory speech representations from unlabelled data. In this work, we propose BRAVEn, an extension to the recent RAVEn method, which learns speech representations entirely from raw audio-visual data. Our modifications to RAVE…

Cited by 0SourceScholar
2024

EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars

CVPR 2024poster

Head avatars animated by visual signals have gained popularity particularly in cross-driving synthesis where the driver differs from the animated character a challenging but highly practical approach. The recently presented MegaPortraits model has demonstrated state-of-the-art results in this domain…

Cited by 26SourcePDFScholar
2024

Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs

NeurIPS 2024poster

Research in auditory, visual, and audiovisual speech recognition (ASR, VSR, and AVSR, respectively) has traditionally been conducted independently. Even recent self-supervised studies addressing two or all three tasks simultaneously tend to yield separate models, leading to disjoint inference pipeli…

2023

Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels

ICASSP 2023accepted

Audio-visual speech recognition has received a lot of attention due to its robustness against acoustic noise. Recently, the performance of automatic, visual, and audio-visual speech recognition (ASR, VSR, and AV-ASR, respectively) has been substantially improved, mainly due to the use of larger mode…

Cited by 0SourceScholar
2023

Jointly Learning Visual and Auditory Speech Representations from Raw Data

ICLR 2023poster

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets generated by slowly-evolving momentum encoders. Driven by the inherent differen…

2023

LA-VOCE: LOW-SNR Audio-Visual Speech Enhancement Using Neural Vocoders

ICASSP 2023accepted

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker’s lip movements. This approach has been shown to yield improvements over audio-only speech enhancement, particularly for the removal of interferin…

Cited by 0SourceScholar
2023

Learning Cross-Lingual Visual Speech Representations

ICASSP 2023accepted

Cross-lingual self-supervised learning has been a growing research topic in the last few years. However, current works only explored the use of audio signals to create representations. In this work, we study cross-lingual self-supervised visual representation learning. We use the recently-proposed R…

Cited by 0SourceScholar
2023

SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision

CVPR 2023poster

Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video datasets are limited in size. In this paper, for the first time, we study the potential of leveraging synthetic visual…

Cited by 27SourcePDFScholar
2022

Leveraging Real Talking Faces via Self-Supervision for Robust Forgery Detection

CVPR 2022poster

One of the most pressing challenges for the detection of face-manipulated videos is generalising to forgery methods not seen during training while remaining effective under common corruptions such as compression. In this paper, we examine whether we can tackle this issue by harnessing videos of real…

Cited by 149PDFcodeScholar
2021

DINO: A Conditional Energy-Based GAN for Domain Translation

ICLR 2021poster

Domain translation is the process of transforming data from one domain to another while preserving the common semantics. Some of the most popular domain translation systems are based on conditional generative adversarial networks, which use source domain data to drive the generator and as an input t…

2021

Lips Don't Lie: A Generalisable and Robust Approach To Face Forgery Detection

CVPR 2021poster

Although current deep learning-based face forgery detectors achieve impressive performance in constrained scenarios, they are vulnerable to samples created by unseen manipulation methods. Some recent works show improvements in generalisation but rely on cues that are easily corrupted by common post-…

Cited by 515PDFcodeScholar
2021

Towards Practical Lipreading with Distilled and Efficient Models

ICASSP 2021accepted

Lipreading has witnessed a lot of progress due to the resurgence of neural networks. Recent works have placed emphasis on aspects such as improving performance by finding the optimal architecture or improving generalization. However, there is still a significant gap between the current methodologies…

Cited by 0SourceScholar
2020

Dynamic Face Video Segmentation via Reinforcement Learning

CVPR 2020poster

For real-time semantic video segmentation, most recent works utilised a dynamic framework with a key scheduler to make online key/non-key decisions. Some works used a fixed key scheduling policy, while others proposed adaptive key scheduling methods based on heuristic strategies, both of which may l…

Cited by 31PDFScholar
2020

Factorized Higher-Order CNNs With an Application to Spatio-Temporal Emotion Estimation

CVPR 2020poster

Training deep neural networks with spatio-temporal (i.e., 3D) or multidimensional convolutions of higher-order is computationally challenging due to millions of unknown parameters across dozens of layers. To alleviate this, one approach is to apply low-rank tensor decompositions to convolution kerne…

Cited by 109PDFScholar
2020

Learning Differentiable Sparse and Low Rank Networks for Audio-Visual Object Localization

ICASSP 2020accepted

Parsimonious modelling, including sparsity and low rankness, has becomes a cornerstone in modern machine learning and signal processing. However, these modelling techniques have limited capabity to learn from large-scale data, and often require some pre-defined parameters to define their optimizatio…

Cited by 0SourceScholar
2020

Multilinear Latent Conditioning for Generating Unseen Attribute Combinations

ICML 2020poster

Deep generative models rely on their inductive bias to facilitate generalization, especially for problems with high dimensional data, like images. However, empirical studies have shown that variational autoencoders (VAE) and generative adversarial networks (GAN) lack the generalization ability that…

Cited by 17SourcePDFScholar
2020

Speech-Driven Facial Animation Using Polynomial Fusion of Features

ICASSP 2020accepted

Speech-driven facial animation involves using a speech signal to generate realistic videos of talking faces. Recent deep learning approaches to facial synthesis rely on extracting low-dimensional representations and concatenating them, followed by a decoding step of the concatenated vector. This acc…

Cited by 0SourceScholar
2020

Towards Pose-Invariant Lip-Reading

ICASSP 2020accepted

Lip-reading models have been significantly improved recently thanks to powerful deep learning architectures. However, most works focused on frontal or near frontal views of the mouth. As a consequence, lip-reading performance seriously deteriorates in non-frontal mouth views. In this work, we presen…

Cited by 0SourceScholar
2020

Visually Guided Self Supervised Learning of Speech Representations

ICASSP 2020accepted

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very limited work that studies the interaction between the two modaliti…

Cited by 0SourceScholar
2019

T-Net: Parametrizing Fully Convolutional Nets With a Single High-Order Tensor

CVPR 2019poster

Recent findings indicate that over-parametrization, while crucial for successfully training deep neural networks, also introduces large amounts of redundancy. Tensor methods have the potential to efficiently parametrize over-complete representations by leveraging this redundancy. In this paper, we p…

Cited by 93PDFScholar
2018

4DFAB: A Large Scale 4D Database for Facial Expression Analysis and Biometric Applications

CVPR 2018poster

The progress we are currently witnessing in many computer vision applications, including automatic face analysis, would not be made possible without tremendous efforts in collecting and annotating large scale visual databases. To this end, we propose 4DFAB, a new large scale database of dynamic high…

Cited by 137SourcePDFScholar
2018

End-to-End Audiovisual Speech Recognition

ICASSP 2018accepted

Several end-to-end deep learning approaches have been recently presented which extract either audio or visual features from the input images or audio signals and perform speech recognition. However, research on end-to-end audiovisual models is very limited. In this work, we present an end-to-end aud…

Cited by 0SourceScholar
2017

Audio-visual object localization and separation using low-rank and sparsity

ICASSP 2017accepted

The ability to localize visual objects that are associated with an audio source and at the same time seperate the audio signal is a corner stone in several audio-visual signal processing applications. Past efforts usually focused on localizing only the visual objects, without audio separation abilit…

Cited by 0SourceScholar
2017

Deep Structured Learning for Facial Action Unit Intensity Estimation

CVPR 2017poster

We consider the task of automated estimation of facial expression intensity. This involves estimation of multiple output variables (facial action units --- AUs) that are structurally dependent. Their structure arises from statistically induced co-occurrence patterns of AU intensity levels. Modeling…

Cited by 155PDFScholar
2017

DeepCoder: Semi-Parametric Variational Autoencoders for Automatic Facial Action Coding

ICCV 2017poster

Human face exhibits an inherent hierarchy in its representations (i.e., holistic facial expressions can be encoded via a set of facial action units (AUs) and their intensity). Variational (deep) auto-encoders (VAE) have shown great results in unsupervised extraction of hierarchical latent representa…

Cited by 56PDFScholar
2016

Copula Ordinal Regression for Joint Estimation of Facial Action Unit Intensity

CVPR 2016poster

Joint modeling of the intensity of facial action units (AUs) from face images is challenging due to the large number of AUs (30+) and their intensity levels (6). This is in part due to the lack of suitable models that can efficiently handle such a large number of outputs/classes simultaneously, but…

Cited by 72PDFScholar
2016

Joint Unsupervised Deformable Spatio-Temporal Alignment of Sequences

CVPR 2016poster

Typically, the problems of spatial and temporal alignment of sequences are considered disjoint. That is, in order to align two sequences, a methodology that (non)-rigidly aligns the images is first applied, followed by temporal alignment of the obtained aligned images. In this paper, we propose the…

Cited by 6PDFScholar
2015

Multi-Conditional Latent Variable Model for Joint Facial Action Unit Detection

ICCV 2015poster

We propose a novel multi-conditional latent variable model for simultaneous facial feature fusion and detection of facial action units. In our approach we exploit the structure-discovery capabilities of generative models such as Gaussian processes, and the discriminative power of classifiers such as…

Cited by 115PDFScholar