← Search

Koki Nagano

18 accepted papers

2025

BLADE: Single-view Body Mesh Estimation through Accurate Depth Estimation

CVPR 2025poster

Single-image human mesh recovery is a challenging task due to the ill-posed nature of simultaneous body shape, pose, and camera estimation. Existing estimators work well on images taken from afar, but they break down as the person moves close to the camera. Moreover, current methods fail to achieve…

Cited by 0SourcePDFScholar
2025

Coherent 3D Portrait Video Reconstruction via Triplane Fusion

CVPR 2025poster

Recent breakthroughs in single-image 3D portrait reconstruction have enabled telepresence systems to stream 3D portrait videos from a single camera in real-time, democratizing telepresence. However, per-frame 3D reconstruction exhibits temporal inconsistency and forgets the user's appearance. On the…

Cited by 1SourcePDFScholar
2025

GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion

ICCV 2025poster

Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture fine-grained dynamic details. To address these limitations,…

Cited by 0SourcePDFScholar
2025

Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation

NeurIPS 2025poster

Synthetic video generation is progressing very rapidly. The latest models can produce very realistic high-resolution videos that are virtually indistinguishable from real ones. Although several video forensic detectors have been recently proposed, they often exhibit poor generalization, which limits…

Cited by 0SourceScholar
2025

Unmasking Puppeteers: Leveraging Biometric Leakage to Expose Impersonation in AI-Based Videoconferencing

NeurIPS 2025poster

AI-based talking-head videoconferencing systems reduce bandwidth by transmitting a latent representation of a speaker’s pose and expression, which is used to synthesize frames on the receiver's end. However, these systems are vulnerable to “puppeteering” attacks, where an adversary controls the iden…

Cited by 0SourceScholar
2024

A Unified Approach for Text- and Image-guided 4D Scene Generation

CVPR 2024poster

Large-scale diffusion generative models are greatly simplifying image video and 3D asset creation from user provided text prompts and images. However the challenging problem of text-to-4D dynamic 3D scene generation with diffusion guidance remains largely unexplored. We propose Dream-in-4D which fea…

Cited by 51SourcePDFScholar
2024

Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos

ECCV 2024poster

"Modern avatar generators allow anyone to synthesize photorealistic real-time talking avatars, ushering in a new era of avatar-based human communication, such as with immersive AR/VR interactions or videoconferencing with limited bandwidths. Their safe adoption, however, requires a mechanism to veri…

Cited by 3SourcePDFScholar
2024

GAvatar: Animatable 3D Gaussian Avatars with Implicit Mesh Learning

CVPR 2024highlight

Gaussian splatting has emerged as a powerful 3D representation that harnesses the advantages of both explicit (mesh) and implicit (NeRF) 3D representations. In this paper we seek to leverage Gaussian splatting to generate realistic animatable avatars from textual descriptions addressing the limitati…

Cited by 41SourcePDFScholar
2024

What You See is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANs

CVPR 2024poster

3D-aware Generative Adversarial Networks (GANs) have shown remarkable progress in learning to generate multi-view-consistent images and 3D geometries of scenes from collections of 2D images via neural volume rendering. Yet the significant memory and computational costs of dense sampling in volume re…

Cited by 8SourcePDFScholar
2023

Generalizable One-shot 3D Neural Head Avatar

NeurIPS 2023poster

We present a method that reconstructs and animates a 3D head avatar from a single-view portrait image. Existing methods either involve time-consuming optimization for a specific person with multiple images, or they struggle to synthesize intricate appearance details beyond the facial region. To addr…

Cited by 31SourcePDFScholar
2023

Generative Novel View Synthesis with 3D-Aware Diffusion Models

ICCV 2023oral

We present a diffusion-based model for 3D-aware generative novel view synthesis from as few as a single input image. Our model samples from the distribution of possible renderings consistent with the input and, even in the presence of ambiguity, is capable of rendering diverse and plausible novel vi…

Cited by 235PDFcodeScholar
2023

On The Detection of Synthetic Images Generated by Diffusion Models

ICASSP 2023accepted

Over the past decade, there has been tremendous progress in creating synthetic media, mainly thanks to the development of powerful methods based on generative adversarial networks (GAN). Very recently, methods based on diffusion models (DM) have been gaining the spotlight. In addition to providing a…

Cited by 0SourceScholar
2023

RANA: Relightable Articulated Neural Avatars

ICCV 2023poster

We propose RANA, a relightable and articulated neural avatar for the photorealistic synthesis of humans under arbitrary viewpoints, body poses, and lighting. We only require a short video clip of the person to create the avatar and assume no knowledge about the lighting environment. We present a nov…

Cited by 16PDFScholar
2022

Efficient Geometry-Aware 3D Generative Adversarial Networks

CVPR 2022oral

Unsupervised generation of high-quality multi-view-consistent images and 3D shapes using only collections of single-view 2D photographs has been a long-standing challenge. Existing 3D GANs are either compute-intensive or make approximations that are not 3D-consistent; the former limits quality and r…

Cited by 1564PDFcodeScholar
2021

Normalized Avatar Synthesis Using StyleGAN and Perceptual Refinement

CVPR 2021poster

We introduce a highly robust GAN-based framework for digitizing a normalized 3D avatar of a person from a single unconstrained photo. While the input image can be of a smiling person or taken in extreme lighting conditions, our method can reliably produce a high-quality textured model of a person's…

Cited by 73PDFScholar
2018

Mesoscopic Facial Geometry Inference Using Deep Neural Networks

CVPR 2018poster

We present a learning-based approach for synthesizing facial geometry at medium and fine scales from diffusely-lit facial texture maps. When applied to an image sequence, the synthesized detail is temporally coherent. Unlike current state-of-the-art methods, which assume "dark is deep", our model…

Cited by 76SourcePDFScholar
2017

Photorealistic Facial Texture Inference Using Deep Neural Networks

CVPR 2017spotlight

We present a data-driven inference method that can synthesize a photorealistic texture map of a complete 3D face model given a partial 2D view of a person in the wild. After an initial estimation of shape and low-frequency albedo, we compute a high-frequency partial texture map, without the shading…

Cited by 162PDFcodeScholar