← Search

Stefanos Zafeiriou

84 accepted papers

2026

Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

CVPR 2026

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of gesture-rich data and their limited ability to infer fine-grained

Cited by 0SourceScholar
2026

Geometric Neural Distance Fields for Learning Human Motion Priors

CVPR 2026

We introduce Neural Riemannian Motion Fields (\name), a novel 3D generative human motion prior that enables robust, temporally consistent, and physically plausible 3D motion recovery. Unlike existing VAE or diffusion-based methods, our higher-order motion prior explicitly models the human motion in

Cited by 0SourceScholar
2026

Parallelised Differentiable Straightest Geodesics for 3D Meshes

CVPR 2026

Machine learning has been progressively generalised to operate within non-Euclidean domains, but geometrically accurate methods for learning on surfaces are still falling behind. The lack of closed-form Riemannian operators, the non-differentiability of their discrete counterparts, and poor parallel

Cited by 0SourceScholar
2026

ReasonX: MLLM-Guided Intrinsic Image Decomposition

CVPR 2026

Intrinsic image decomposition aims to separate images into physical components such as albedo, depth, normals, and illumination. While recent diffusion- and transformer-based models benefit from paired supervision from synthetic datasets, their generalization to diverse, real-world scenarios remains

Cited by 0SourcecodeScholar
2026

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

ICML 2026poster

This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and free-viewpoint talking portrait synthesis, given an identity embedding or reference image, within a unified framework. Existing 2D speech-to-video diffusio…

Cited by 0SourceScholar
2026

SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action Planning

CVPR 2026

Vision-Language-Action (VLA) models have emerged as a promising paradigm where pretrained Vision-Language Models (VLMs) serve as System 2 for high-level reasoning, connected to action experts as System 1 for low-level motor control.However, current works fail to genuinely leverage VLM capabilities:

Cited by 0SourceScholar
2025

Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidance

CVPR 2025poster

Inspired by the effectiveness of 3D Gaussian Splatting (3DGS) in reconstructing detailed 3D scenes within multi-view setups and the emergence of large 2D human foundation models, we introduce Arc2Avatar, the first SDS-based method utilizing a human face foundation model as guidance with just a singl…

Cited by 2SourcePDFScholar
2025

Are Large Brainwave Foundation Models Capable Yet ? Insights from Fine-Tuning

ICML 2025poster

Foundation Models have demonstrated significant success across various domains in Artificial Intelligence (AI), yet their capabilities for brainwave modeling remain unclear. In this paper, we comprehensively evaluate current Large Brainwave Foundation Models (LBMs) through systematic fine-tuning exp…

Cited by 0SourcePDFScholar
2025

Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera

CVPR 2025highlight

We propose Dyn-HaMR, to the best of our knowledge, the first approach to reconstruct 4D global hand motion from monocular videos recorded by dynamic cameras in the wild. Reconstructing accurate 3D hand meshes from monocular videos is a crucial task for understanding human behaviour, with significant…

Cited by 1SourcePDFScholar
2025

ImHead: A Large-scale Implicit Morphable Model for Localized Head Modeling

ICCV 2025poster

Over the last years, 3D morphable models (3DMMs) have emerged as a state-of-the-art methodology for modeling and generating expressive 3D avatars. However, given their reliance on a strict topology, along with their linear nature, they struggle to represent complex full-head shapes. Following the ad…

Cited by 0SourcePDFScholar
2025

Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility

ICCV 2025poster

Robustness and resource-efficiency are two highly desirable properties for modern machine learning models. However, achieving them jointly remains a challenge. In this paper, we identify high learning rates as a facilitator for simultaneously achieving robustness to spurious correlations and network…

2025

S^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion Priors

CVPR 2025poster

Recent 3D face reconstruction methods have made remarkable advancements, yet achieving high-quality facial reflectance from monocular input remains challenging. Existing methods rely on the light-stage captured data to learn facial reflectance models. However, limited subject diversity in these data…

Cited by 0SourcePDFScholar
2025

Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator

ICCV 2025poster

Sign language is a visual language that encompasses all linguistic features of natural languages and serves as the primary communication method for the deaf and hard-of-hearing communities. Although many studies have successfully adapted pretrained language models (LMs) for sign language translation…

2025

SpinMeRound: Consistent Multi-View Identity Generation Using Diffusion Models

ICCV 2025poster

Despite recent progress in diffusion models, generating realistic head portraits from novel viewpoints remains a significant challenge in computer vision. Most current approaches are constrained to limited angular ranges, predominantly focusing on frontal or near-frontal views. Moreover, although th…

Cited by 0SourcePDFScholar
2025

WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild

CVPR 2025poster

In recent years, 3D hand pose estimation methods have garnered significant attention due to their extensive applications in human-computer interaction, virtual reality, and robotics. In contrast, there has been a notable gap in hand detection pipelines, posing significant challenges in constructing…

2024

3DGazeNet: Generalizing Gaze Estimation with Weak Supervision from Synthetic Views

ECCV 2024poster

"Developing gaze estimation models that generalize well to unseen domains and in-the-wild conditions remains a challenge with no known best solution. This is mostly due to the difficulty of acquiring ground truth data that cover the distribution of faces, head poses, and environments that exist in t…

Cited by 8SourcePDFScholar
2024

AnimateMe: 4D Facial Expressions via Diffusion Models

ECCV 2024poster

"The field of photorealistic 3D avatar reconstruction and generation has garnered significant attention in recent years; however, animating such avatars remains challenging. Recent advances in diffusion models have notably enhanced the capabilities of generative models in 2D animation. In this work,…

Cited by 2SourcePDFScholar
2024

Arc2Face: A Foundation Model for ID-Consistent Human Faces

ECCV 2024oral

"This paper presents , an identity-conditioned face foundation model, which, given the ArcFace embedding of a person, can generate diverse photo-realistic images with an unparalleled degree of face similarity than existing models. Despite previous attempts to decode face recognition features into de…

2024

Distribution Matching for Multi-Task Learning of Classification Tasks: A Large-Scale Study on Faces & Beyond

AAAI 2024technical

Multi-Task Learning (MTL) is a framework, where multiple related tasks are learned jointly and benefit from a shared representation space, or parameter transfer. To provide sufficient learning support, modern MTL uses annotated data with full, or sufficiently large overlap across tasks, i.e., each i…

Cited by 39SourcePDFScholar
2024

ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling

NeurIPS 2024poster

We propose ID-to-3D, a method to generate identity- and text-guided 3D human heads with disentangled expressions, starting from even a single casually captured ‘in-the-wild’ image of a subject. The foundation of our approach is anchored in compositionality, alongside the use of task-specific 2D diff…

Cited by 2SourcePDFScholar
2024

Locally Adaptive Neural 3D Morphable Models

CVPR 2024poster

We present the Locally Adaptive Morphable Model (LAMM) a highly flexible Auto-Encoder (AE) framework for learning to generate and manipulate 3D meshes. We train our architecture following a simple self-supervised training scheme in which input displacements over a set of sparse control vertices are…

2024

Neural Sign Actors: A Diffusion Model for 3D Sign Language Production from Text

CVPR 2024poster

Sign Languages (SL) serve as the primary mode of communication for the Deaf and Hard of Hearing communities. Deep learning methods for SL recognition and translation have achieved promising results. However Sign Language Production (SLP) poses a challenge as the generated motions must be realistic a…

Cited by 20SourcePDFScholar
2024

SAGS: Structure-Aware 3D Gaussian Splatting

ECCV 2024poster

"Following the advent of NeRFs, 3D Gaussian Splatting (3D-GS) has paved the way to real-time neural rendering overcoming the computational burden of volumetric methods. Several extensions of 3D-GS have been proposed to achieve compressible and high-fidelity performance. However, by employing a geome…

Cited by 7SourcePDFScholar
2024

Shapefusion: 3D localized human diffusion models

ECCV 2024poster

"In the realm of 3D computer vision, parametric models have emerged as a ground-breaking methodology for the creation of realistic and expressive 3D avatars. Traditionally, they rely on Principal Component Analysis (PCA), given its ability to decompose data to an orthonormal space that maximally cap…

2024

UV-free Texture Generation with Denoising and Geodesic Heat Diffusion

NeurIPS 2024poster

Seams, distortions, wasted UV space, vertex-duplication, and varying resolution over the surface are the most prominent issues of the standard UV-based texturing of meshes. These issues are particularly acute when automatic UV-unwrapping techniques are used. For this reason, instead of generating te…

2023

Adaptive Spiral Layers for Efficient 3D Representation Learning on Meshes

ICCV 2023poster

The success of deep learning models on structured data has generated significant interest in extending their application to non-Euclidean domains. In this work, we introduce a novel intrinsic operator suitable for representation learning on 3D meshes. Our operator is specifically tailored to adapt i…

Cited by 0PDFcodeScholar
2023

FitMe: Deep Photorealistic 3D Morphable Model Avatars

CVPR 2023poster

In this paper, we introduce FitMe, a facial reflectance model and a differentiable rendering optimization pipeline, that can be used to acquire high-fidelity renderable human avatars from single or multiple images. The model consists of a multi-modal style-based generator, that captures facial appea…

Cited by 35SourcePDFScholar
2023

Handy: Towards a High Fidelity 3D Hand Shape and Appearance Model

CVPR 2023poster

Over the last few years, with the advent of virtual and augmented reality, an enormous amount of research has been focused on modeling, tracking and reconstructing human hands. Given their power to express human behavior, hands have been a very important, but challenging component of the human body.…

2023

Relightify: Relightable 3D Faces from a Single Image via Diffusion Models

ICCV 2023poster

Following the remarkable success of diffusion models on image generation, recent works have also demonstrated their impressive ability to address a number of inverse problems in an unsupervised way, by properly constraining the sampling process based on a conditioning input. Motivated by this, in th…

Cited by 28PDFcodeScholar
2023

Spatio-temporal Prompting Network for Robust Video Feature Extraction

ICCV 2023poster

The frame quality deterioration problem is one of the main challenges in the field of video understanding. To compensate for the information loss caused by deteriorated frames, recent approaches exploit transformer-based integration modules to obtain spatio-temporal information. However, these integ…

Cited by 5PDFcodeScholar
2023

ViTs for SITS: Vision Transformers for Satellite Image Time Series

CVPR 2023poster

In this paper we introduce the Temporo-Spatial Vision Transformer (TSViT), a fully-attentional model for general Satellite Image Time Series (SITS) processing based on the Vision Transformer (ViT). TSViT splits a SITS record into non-overlapping patches in space and time which are tokenized and subs…

2022

3D Human Tongue Reconstruction From Single "In-the-Wild" Images

CVPR 2022oral

3D face reconstruction from a single image is a task that has garnered increased interest in the Computer Vision community, especially due to its broad use in a number of applications such as realistic 3D avatar creation, pose invariant face recognition and face hallucination. Since the introduction…

Cited by 7PDFcodeScholar
2022

Decoupled Multi-Task Learning With Cyclical Self-Regulation for Face Parsing

CVPR 2022poster

This paper probes intrinsic factors behind typical failure cases (e.g spatial inconsistency and boundary confusion) produced by the existing state-of-the-art method in face parsing. To tackle these problems, we propose a novel Decoupled Multi-task Learning with Cyclical Self-Regulation (DML-CSR) for…

Cited by 43PDFcodeScholar
2022

MimicME: A Large Scale Diverse 4D Database for Facial Expression Analysis

ECCV 2022poster

"Recently, Deep Neural Networks (DNNs) have been shown to outperform traditional methods in many disciplines such as computer vision, speech recognition and natural language processing. A prerequisite for the successful application of DNNs is the big number of data. Even though various facial datase…

2022

Physically-Based Face Rendering for NIR-VIS Face Recognition

NeurIPS 2022accept

Near infrared (NIR) to Visible (VIS) face matching is challenging due to the significant domain gaps as well as a lack of sufficient data for cross-modality model training. To overcome this problem, we propose a novel method for paired NIR-VIS facial image generation. Specifically, we reconstruct 3D…

2022

Revisiting Point Cloud Simplification: A Learnable Feature Preserving Approach

ECCV 2022poster

"The recent advances in 3D sensing technology have made possible the capture of point clouds in significantly high resolution. However, increased detail usually comes at the expense of high storage, as well as computational costs in terms of processing and visualization operations. Mesh and Point Cl…

Cited by 34SourcePDFScholar
2022

Sample and Computation Redistribution for Efficient Face Detection

ICLR 2022poster

Although tremendous strides have been made in uncontrolled face detection, accurate face detection with a low computation cost remains an open challenge. In this paper, we point out that computation distribution and scale augmentation are the keys to detecting small faces from low-resolution images.…

2021

Poly-NL: Linear Complexity Non-Local Layers With 3rd Order Polynomials

ICCV 2021poster

Spatial self-attention layers, in the form of Non-Local blocks, introduce long-range dependencies in Convolutional Neural Networks by computing pairwise similarities among all possible positions. Such pairwise functions underpin the effectiveness of non-local layers, but also determine a complexity…

Cited by 14PDFScholar
2021

Speech Emotion Recognition Using Semantic Information

ICASSP 2021accepted

Speech emotion recognition is a crucial problem manifesting in a multitude of applications such as human computer interaction and education. Although several advancements have been made in the recent years, especially with the advent of Deep Neural Networks (DNN), most of the studies in the literatu…

Cited by 0SourceScholar
2021

Variational Prototype Learning for Deep Face Recognition

CVPR 2021poster

Deep face recognition has achieved remarkable improvements due to the introduction of margin-based softmax loss, in which the prototype stored in the last linear layer represents the center of each class. In these methods, training samples are enforced to be close to positive prototypes and far apar…

Cited by 100PDFScholar
2020

AvatarMe: Realistically Renderable 3D Facial Reconstruction "In-the-Wild"

CVPR 2020poster

Over the last years, with the advent of Generative Adversarial Networks (GANs), many face analysis tasks have accomplished astounding performance, with applications including, but not limited to, face generation and 3D face reconstruction from a single "in-the-wild" image. Nevertheless, to the best…

Cited by 201PDFcodeScholar
2020

BLSM: A Bone-Level Skinned Model of the Human Mesh

ECCV 2020poster

We introduce BLSM, a bone-level skinned model of the human body mesh where bone scales are set prior to template synthesis, rather than the common, inverse practice. BLSM first sets bone lengths and joint angles to specify the skeleton, then specifies identity-specific surface variation, and finally…

Cited by 21SourcePDFScholar
2020

DeepFaceFlow: In-the-Wild Dense 3D Facial Motion Estimation

CVPR 2020poster

Dense 3D facial motion capture from only monocular in-the-wild pairs of RGB images is a highly challenging problem with numerous applications, ranging from facial expression recognition to facial reenactment. In this work, we propose DeepFaceFlow, a robust, fast, and highly-accurate framework for th…

Cited by 17PDFcodeScholar
2020

Geometrically Principled Connections in Graph Neural Networks

CVPR 2020poster

Graph convolution operators bring the advantages of deep learning to a variety of graph and mesh processing tasks previously deemed out of reach. With their continued success comes the desire to design more powerful architectures, often by adapting existing deep learning techniques to non-Euclidean…

Cited by 29PDFScholar
2020

Learning to Generate Customized Dynamic 3D Facial Expressions

ECCV 2020poster

Recent advances in deep learning have significantly pushed the state-of-the-art in photorealistic video animation given a single image. In this paper, we extrapolate those advances to the 3D domain, by studying 3D image-to-video translation with a particular focus on 4D facial expressions. Although…

Cited by 25SourcePDFScholar
2020

P-nets: Deep Polynomial Neural Networks

CVPR 2020poster

Deep Convolutional Neural Networks (DCNNs) is currently the method of choice both for generative, as well as for discriminative learning in computer vision and machine learning. The success of DCNNs can be attributed to the careful selection of their building blocks (e.g., residual blocks, rectifier…

Cited by 95PDFcodeScholar
2020

Reconstructing the Noise Variance Manifold for Image Denoising

ECCV 2020poster

Deep Convolutional Neural Networks (CNNs) have been successfully used in many low-level vision problems like image denoising. Although the conditional image generation techniques have led to large improvements in this task, there has been little effort in providing conditional generative adversarial…

Cited by 8SourcePDFScholar
2020

RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild

CVPR 2020poster

Though tremendous strides have been made in uncontrolled face detection, accurate and efficient 2D face alignment and 3D face reconstruction in-the-wild remain an open challenge. In this paper, we present a novel single-shot, multi-level face localisation method, named RetinaFace, which unifies face…

Cited by 1569PDFScholar
2020

Sub-center ArcFace: Boosting Face Recognition by Large-scale Noisy Web Faces

ECCV 2020poster

Margin-based deep face recognition methods (e.g. SphereFace, CosFace, and ArcFace) have achieved remarkable success in unconstrained face recognition. However, these methods are susceptible to the massive label noise in the training data and thus require laborious human effort to clean the datasets.…

2020

Synthesizing Coupled 3D Face Modalities by Trunk-Branch Generative Adversarial Networks

ECCV 2020poster

Generating realistic 3D faces is of high importance for computer graphics and computer vision applications. Generally, research on 3D face generation revolves around linear statistical models of the facial surface. Nevertheless, these models cannot represent faithfully either the facial texture or t…

2020

TESA: Tensor Element Self-Attention via Matricization

CVPR 2020poster

Representation learning is a fundamental part of modern computer vision, where abstract representations of data are encoded as tensors optimized to solve problems like image segmentation and inpainting. Recently, self-attention in the form of Non-Local Block has emerged as a powerful technique to en…

Cited by 26PDFcodeScholar
2020

Weakly-Supervised Mesh-Convolutional Hand Reconstruction in the Wild

CVPR 2020oral

We introduce a simple and effective network architecture for monocular 3D hand pose estimation consisting of an image encoder followed by a mesh convolutional decoder that is trained through a direct 3D hand mesh reconstruction loss. We train our network by gathering a large-scale dataset of hand ac…

Cited by 247PDFScholar
2019

ArcFace: Additive Angular Margin Loss for Deep Face Recognition

CVPR 2019oral

One of the main challenges in feature learning using Deep Convolutional Neural Networks (DCNNs) for large-scale face recognition is the design of appropriate loss functions that can enhance the discriminative power. Centre loss penalises the distance between deep features and their corresponding cla…

Cited by 8474PDFcodeScholar
2019

Combining 3D Morphable Models: A Large Scale Face-And-Head Model

CVPR 2019oral

Three-dimensional Morphable Models (3DMMs) are powerful statistical tools for representing the 3D surfaces of an object class. In this context, we identify an interesting question that has previously not received research attention: is it possible to combine two or more 3DMMs that (a) are built usin…

Cited by 104PDFScholar
2019

Dense 3D Face Decoding Over 2500FPS: Joint Texture & Shape Convolutional Mesh Decoders

CVPR 2019poster

3D Morphable Models (3DMMs) are statistical models that represent facial texture and shape variations using a set of linear bases and more particular Principal Component Analysis (PCA). 3DMMs were used as statistical priors for reconstructing 3D faces from images by solving non-linear least square o…

Cited by 93PDFScholar
2019

GANFIT: Generative Adversarial Network Fitting for High Fidelity 3D Face Reconstruction

CVPR 2019poster

In the past few years, a lot of work has been done towards reconstructing the 3D facial structure from single images by capitalizing on the power of Deep Convolutional Neural Networks (DCNNs). In the most recent works, differentiable renderers were employed in order to learn the relationship between…

Cited by 414PDFcodeScholar
2019

Neural 3D Morphable Models: Spiral Convolutional Networks for 3D Shape Representation Learning and Generation

ICCV 2019poster

Generative models for 3D geometric data arise in many important applications in 3D computer vision and graphics. In this paper, we focus on 3D deformable shapes that share a common topological structure, such as human faces and bodies. Morphable Models and their variants, despite their linear formul…

Cited by 192PDFcodeScholar
2019

Robust Conditional Generative Adversarial Networks

ICLR 2019poster

Conditional generative adversarial networks (cGAN) have led to large improvements in the task of conditional image generation, which lies at the heart of computer vision. The major focus so far has been on performance improvement, while there has been little effort in making cGAN more robust to nois…

2018

4DFAB: A Large Scale 4D Database for Facial Expression Analysis and Biometric Applications

CVPR 2018poster

The progress we are currently witnessing in many computer vision applications, including automatic face analysis, would not be made possible without tremendous efforts in collecting and annotating large scale visual databases. To this end, we propose 4DFAB, a new large scale database of dynamic high…

Cited by 137SourcePDFScholar
2018

UV-GAN: Adversarial Facial UV Map Completion for Pose-Invariant Face Recognition

CVPR 2018poster

Recently proposed robust 3D face alignment methods establish either dense or sparse correspondence between a 3D face model and a 2D facial image. The use of these methods presents new challenges as well as opportunities for facial texture analysis. In particular, by sampling the image using the fitt…

2017

3D Face Morphable Models "In-The-Wild"

CVPR 2017spotlight

3D Morphable Models (3DMMs) are powerful statistical models of 3D facial shape and texture, and among the state-of-the-art methods for reconstructing facial shape from single images. With the advent of new 3D sensors, many 3D facial datasets have been collected containing both neutral as well as exp…

Cited by 213PDFScholar
2017

DenseReg: Fully Convolutional Dense Shape Regression In-The-Wild

CVPR 2017poster

In this paper we propose to learn a mapping from image pixels into a dense template grid through a fully convolutional network. We formulate this task as a regression problem and train our network by leveraging upon manually annotated facial landmarks 'in-the-wild'. We use such landmarks to establ…

Cited by 239PDFScholar
2017

Dynamic Probabilistic Linear Discriminant Analysis for video classification

ICASSP 2017accepted

Component Analysis (CA) comprises of statistical techniques that decompose signals into appropriate latent components, relevant to a task-at-hand (e.g., clustering, segmentation, classification). Recently, an explosion of research in CA has been witnessed, with several novel probabilistic models pro…

Cited by 0SourceScholar
2017

Face Normals "In-The-Wild" Using Fully Convolutional Networks

CVPR 2017poster

In this work we pursue a data-driven approach to the problem of estimating surface normals from a single intensity image, focusing in particular on human faces. We introduce new methods to exploit the currently available facial databases for dataset construction and tailor a deep convolutional neura…

Cited by 58PDFScholar
2017

Robust Kronecker-Decomposable Component Analysis for Low-Rank Modeling

ICCV 2017poster

Dictionary learning and component analysis are part of one of the most well-studied and active research fields, at the intersection of signal and image processing, computer vision, and statistical machine learning. In dictionary learning, the current methods of choice are arguably K-SVD and its vari…

Cited by 14PDFcodeScholar
2017

Side Information in Robust Principal Component Analysis: Algorithms and Applications

ICCV 2017poster

Robust Principal Component Analysis (RPCA) aims at recovering a low-rank subspace from grossly corrupted high-dimensional (often visual) data and is a cornerstone in many machine learning and computer vision applications. Even though RPCA has been shown to be very successful in solving many rank min…

Cited by 12PDFScholar
2016

A 3D Morphable Model Learnt From 10,000 Faces

CVPR 2016spotlight

We present Large Scale Facial Model (LSFM) -- a 3D Morphable Model (3DMM) automatically constructed from 9,663 distinct facial identities. To the best of our knowledge LSFM is the largest-scale Morphable Model ever constructed, containing statistical information from a huge variety of the human popu…

Cited by 424PDFScholar
2016

Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent network

ICASSP 2016accepted

The automatic recognition of spontaneous emotions from speech is a challenging task. On the one hand, acoustic features need to be robust enough to capture the emotional content for various styles of speaking, and while on the other, machine learning algorithms need to be insensitive to outliers whi…

Cited by 0SourceScholar
2016

Estimating Correspondences of Deformable Objects "In-The-Wild"

CVPR 2016poster

During the past few years we have witnessed the development of many methodologies for building and fitting Statistical Deformable Models (SDMs). The construction of accurate SDMs requires careful annotation of images with regards to a consistent set of landmarks. However, the manual annotation of a…

Cited by 12PDFScholar
2016

Joint Unsupervised Deformable Spatio-Temporal Alignment of Sequences

CVPR 2016poster

Typically, the problems of spatial and temporal alignment of sequences are considered disjoint. That is, in order to align two sequences, a methodology that (non)-rigidly aligns the images is first applied, followed by temporal alignment of the obtained aligned images. In this paper, we propose the…

Cited by 6PDFScholar
2016

Mnemonic Descent Method: A Recurrent Process Applied for End-To-End Face Alignment

CVPR 2016poster

Cascaded regression has recently become the method of choice for solving non-linear least squares problems such as deformable image alignment. Given a sizeable training set, cascaded regression learns a set of generic rules that are sequentially applied to minimise the least squares problem. Despit…

Cited by 444PDFScholar
2015

Automatic Construction Of Robust Spherical Harmonic Subspaces

CVPR 2015poster

In this paper we propose a method to automatically recover a class specific low dimensional spherical harmonic basis from a set of in-the-wild facial images. We combine existing techniques for uncalibrated photometric stereo and low rank matrix decompositions in order to robustly recover a combined…

Cited by 23SourcePDFScholar