← Search

Sean Fanello

17 accepted papers

2026

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

CVPR 2026

Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open challenge. In this paper, we present Archon, a fully pretrained, human-centric unified multimodal model for holistic avata

Cited by 0SourceScholar
2026

Talking Together: Synthesizing Co-Located 3D Conversations from Audio

CVPR 2026

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference call, our work is the first to explicitly model the dynamic 3

Cited by 0SourceScholar
2025

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular VideosC

CVPR 2025poster

We propose a novel 3D-aware diffusion-based method for generating photorealistic talking head videos directly from a single identity image and explicit control signals (e.g., expressions). Our method generates Multiplane Images (MPIs) that ensure geometric consistency, making them ideal for immersiv…

Cited by 0SourcePDFScholar
2025

SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix

ICLR 2025poster

Video generation models have demonstrated great capability of producing impressive monocular videos, however, the generation of 3D stereoscopic video remains under-explored. We propose a pose-free and training-free approach for generating 3D stereoscopic videos using an off-the-shelf monocular video…

2024

Efficient 3D Implicit Head Avatar with Mesh-anchored Hash Table Blendshapes

CVPR 2024poster

3D head avatars built with neural implicit volumetric representations have achieved unprecedented levels of photorealism. However the computational cost of these methods remains a significant barrier to their widespread adoption particularly in real-time applications such as virtual reality and tele…

Cited by 5SourcePDFScholar
2024

Loc3Diff: Local Diffusion for 3D Human Head Synthesis and Editing

ECCV 2024poster

"We present a novel framework for generating photorealistic 3D human head and subsequently manipulating and reposing them with remarkable flexibility. The proposed approach constructs an implicit representation of 3D human heads, anchored on a parametric face model. To enhance representational capab…

Cited by 0SourcePDFScholar
2024

MVDD: Multi-View Depth Diffusion Models

ECCV 2024poster

"Denoising diffusion models have demonstrated outstanding results in 2D image generation, yet it remains a challenge to replicate its success in 3D shape generation. In this paper, we propose leveraging multi-view depth, which represents complex 3D shapes in a 2D data format that is easy to denoise.…

Cited by 4SourcePDFScholar
2023

Controllable Light Diffusion for Portraits

CVPR 2023poster

We introduce light diffusion, a novel method to improve lighting in portraits, softening harsh shadows and specular highlights while preserving overall scene illumination. Inspired by professional photographers' diffusers and scrims, our method softens lighting given only a single portrait photo. Pr…

Cited by 12SourcePDFScholar
2023

Learning Personalized High Quality Volumetric Head Avatars From Monocular RGB Videos

CVPR 2023poster

We propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our hybrid pipeline combines the geometry prior and dynamic tracki…

Cited by 20SourcePDFScholar
2021

HITNet: Hierarchical Iterative Tile Refinement Network for Real-time Stereo Matching

CVPR 2021poster

This paper presents HITNet, a novel neural network architecture for real-time stereo matching. Contrary to many recent neural network approaches that operate on a full costvolume and rely on 3D convolutions, our approach does not explicitly build a volume and instead relies on a fast multi-resolutio…

Cited by 326PDFcodeScholar
2021

HumanGPS: Geodesic PreServing Feature for Dense Human Correspondences

CVPR 2021poster

In this paper, we address the problem of building pixel-wise dense correspondences between human images under arbitrary camera viewpoints and body poses. Previous methods either assume small motions or rely on discriminative descriptors extracted from local patches, which cannot handle large motion…

Cited by 14PDFScholar
2021

Multiresolution Deep Implicit Functions for 3D Shape Representation

ICCV 2021poster

We introduce Multiresolution Deep Implicit Functions (MDIF), a hierarchical representation that can recover fine geometry detail, while being able to perform global operations such as shape completion. Our model represents a complex 3D shape with a hierarchy of latent grids, which can be decoded int…

Cited by 52PDFScholar
2020

Deep Implicit Volume Compression

CVPR 2020oral

We describe a novel approach for compressing truncated signed distance fields (TSDF) stored in 3D voxel grids, and their corresponding textures. To compress the TSDF, our method relies on a block-based neural network architecture trained end-to-end, achieving state-of-the-art rate-distortion trade-o…

Cited by 52PDFcodeScholar
2020

Du²Net: Learning Depth Estimation from Dual-Cameras and Dual-Pixels

ECCV 2020poster

Computational stereo has reached a high level of accuracy, but degrades in the presence of occlusions, repeated textures, and correspondence errors along edges. We present a novel approach based on neural networks for depth estimation that combines stereo from dual cameras with stereo from a dual-pi…

Cited by 39SourcePDFScholar
2019

Volumetric Capture of Humans With a Single RGBD Camera via Semi-Parametric Learning

CVPR 2019poster

Volumetric (4D) performance capture is fundamental for AR/VR content generation. Whereas previous work in 4D performance capture has shown impressive results in studio settings, the technology is still far from being accessible to a typical consumer who, at best, might own a single RGBD sensor. Thus…

Cited by 47PDFScholar
2018

ActiveStereoNet: End-to-End Self-Supervised Learning for Active Stereo Systems

ECCV 2018poster

In this paper we present ActiveStereoNet, the first deep learning solution for active stereo systems. Due to the lack of ground truth, our method is fully self-supervised, yet it produces precise depth with a subpixel precision of 1/30th of a pixel; it does not suffer from the common over-smoothing…

Cited by 139SourcePDFScholar
2018

StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction

ECCV 2018poster

This paper presents StereoNet, the first end-to-end deep architecture for real-time stereo matching that runs at 60 fps on an NVidia Titan X, producing high-quality, edge-preserved, quantization-free depth maps. A key insight of this paper is that the network achieves a sub-pixel matching precision…

Cited by 461SourcePDFScholar