← Search

Konstantinos Vougioukas

7 accepted papers

2025

KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation

CVPR 2025poster

Current audio-driven facial animation methods achieve impressive results for short videos but suffer from error accumulation and identity drift when extended to longer durations. Existing methods attempt to mitigate this through external spatial control, increasing long-term consistency but compromi…

Cited by 2SourcePDFScholar
2024

EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars

CVPR 2024poster

Head avatars animated by visual signals have gained popularity particularly in cross-driving synthesis where the driver differs from the animated character a challenging but highly practical approach. The recently presented MegaPortraits model has demonstrated state-of-the-art results in this domain…

Cited by 26SourcePDFScholar
2023

SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision

CVPR 2023poster

Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video datasets are limited in size. In this paper, for the first time, we study the potential of leveraging synthetic visual…

Cited by 27SourcePDFScholar
2021

DINO: A Conditional Energy-Based GAN for Domain Translation

ICLR 2021poster

Domain translation is the process of transforming data from one domain to another while preserving the common semantics. Some of the most popular domain translation systems are based on conditional generative adversarial networks, which use source domain data to drive the generator and as an input t…

2021

Lips Don't Lie: A Generalisable and Robust Approach To Face Forgery Detection

CVPR 2021poster

Although current deep learning-based face forgery detectors achieve impressive performance in constrained scenarios, they are vulnerable to samples created by unseen manipulation methods. Some recent works show improvements in generalisation but rely on cues that are easily corrupted by common post-…

Cited by 515PDFcodeScholar
2020

Speech-Driven Facial Animation Using Polynomial Fusion of Features

ICASSP 2020accepted

Speech-driven facial animation involves using a speech signal to generate realistic videos of talking faces. Recent deep learning approaches to facial synthesis rely on extracting low-dimensional representations and concatenating them, followed by a decoding step of the concatenated vector. This acc…

Cited by 0SourceScholar
2020

Visually Guided Self Supervised Learning of Speech Representations

ICASSP 2020accepted

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very limited work that studies the interaction between the two modaliti…

Cited by 0SourceScholar