← Search

Iasonas Kokkinos

22 accepted papers

2024

MeshPose: Unifying DensePose and 3D Body Mesh Reconstruction

CVPR 2024poster

DensePose provides a pixel-accurate association of images with 3D mesh coordinates but does not provide a 3D mesh while Human Mesh Reconstruction (HMR) systems have high 2D reprojection error as measured by DensePose localization metrics. In this work we introduce MeshPose to jointly tackle DensePos…

2023

StyleMorph: Disentangled 3D-Aware Image Synthesis with a 3D Morphable StyleGAN

ICLR 2023poster

We introduce StyleMorph, a 3D-aware generative model that disentangles 3D shape, camera pose, object appearance, and background appearance for high quality image synthesis. We account for shape variability by morphing a canonical 3D object template, effectively learning a 3D morphable model in an en…

Cited by 3SourcePDFScholar
2020

BLSM: A Bone-Level Skinned Model of the Human Mesh

ECCV 2020poster

We introduce BLSM, a bone-level skinned model of the human body mesh where bone scales are set prior to template synthesis, rather than the common, inverse practice. BLSM first sets bone lengths and joint angles to specify the skeleton, then specifies identity-specific surface variation, and finally…

Cited by 21SourcePDFScholar
2020

Weakly-Supervised Mesh-Convolutional Hand Reconstruction in the Wild

CVPR 2020oral

We introduce a simple and effective network architecture for monocular 3D hand pose estimation consisting of an image encoder followed by a mesh convolutional decoder that is trained through a direct 3D hand mesh reconstruction loss. We train our network by gathering a large-scale dataset of hand ac…

Cited by 247PDFScholar
2019

Slim DensePose: Thrifty Learning From Sparse Annotations and Motion Cues

CVPR 2019oral

DensePose supersedes traditional landmark detectors by densely mapping image pixels to body surface coordinates. This power, however, comes at a greatly increased annotation cost, as supervising the model requires to manually label hundreds of points per pose instance. In this work, we thus seek met…

Cited by 40PDFScholar
2018

Deep Spatio-Temporal Random Fields for Efficient Video Segmentation

CVPR 2018poster

In this work we introduce a time- and memory-efficient method for structured prediction that couples neuron decisions across both space at time. We show that we are able to perform exact and efficient inference on a densely connected spatio-temporal graph by capitalizing on recent advances on deep…

2018

Deforming Autoencoders: Unsupervised Disentangling of Shape and Appearance

ECCV 2018poster

In this work we introduce the Deforming Autoencoder, a generative model for images that disentangles shape from appearance in a latent representation space that is learned in a fully unsupervised manner. As in the deformable template paradigm, shape is represented as a diffeomorphism between a canon…

Cited by 249SourcePDFScholar
2018

Learning Filterbanks from Raw Speech for Phone Recognition

ICASSP 2018accepted

We train a bank of complex filters that operates on the raw waveform and is fed into a convolutional neural network for end-to-end phone recognition. These time-domain filterbanks (TD-filterbanks) are initialized as an approximation of mel-filterbanks, and then fine-tuned jointly with the remaining…

Cited by 0SourceScholar
2017

DenseReg: Fully Convolutional Dense Shape Regression In-The-Wild

CVPR 2017poster

In this paper we propose to learn a mapping from image pixels into a dense template grid through a fully convolutional network. We formulate this task as a regression problem and train our network by leveraging upon manually annotated facial landmarks 'in-the-wild'. We use such landmarks to establ…

Cited by 239PDFScholar
2017

Face Normals "In-The-Wild" Using Fully Convolutional Networks

CVPR 2017poster

In this work we pursue a data-driven approach to the problem of estimating surface normals from a single intensity image, focusing in particular on human faces. We introduce new methods to exploit the currently available facial databases for dataset construction and tailor a deep convolutional neura…

Cited by 58PDFScholar
2017

Segmentation-Aware Convolutional Networks Using Local Attention Masks

ICCV 2017poster

We introduce an approach to integrate segmentation information within a convolutional neural network (CNN). This counter-acts the tendency of CNNs to smooth information across regions and increases their spatial precision. To obtain segmentation information, we set up a CNN to provide an embedding s…

Cited by 188PDFcodeScholar
2017

Ubernet: Training a Universal Convolutional Neural Network for Low-, Mid-, and High-Level Vision Using Diverse Datasets and Limited Memory

CVPR 2017oral

In this work we train in an end-to-end manner a convolutional neural network (CNN) that jointly handles low-, mid-, and high-level vision tasks in a unified architecture. Such a network can act like a `swiss knife' for vision tasks; we call it an "UberNet" to indicate its overarching nature.…

Cited by 855PDFScholar
2015

Discriminative Learning of Deep Convolutional Feature Point Descriptors

ICCV 2015poster

Deep learning has revolutionalized image-level tasks such as classification, but patch-level tasks, such as correspondence, still rely on hand-crafted features, e.g. SIFT. In this paper we use Convolutional Neural Networks (CNNs) to learn discriminant patch representations and in particular train a…

Cited by 1022PDFcodeScholar
2015

Modeling Local and Global Deformations in Deep Learning: Epitomic Convolution, Multiple Instance Learning, and Sliding Window Detection

CVPR 2015poster

Deep Convolutional Neural Networks (DCNNs) achieve invariance to domain transformations (deformations) by using multiple 'max-pooling' (MP) layers. In this work we show that alternative methods of modeling deformations can improve the accuracy and efficiency of DCNNs. First, we introduce epitomic co…

Cited by 252SourcePDFScholar