← Search

Rolandos Alexandros Potamias

21 accepted papers

2026

CEDex: Cross-Embodiment Dexterous Grasp Generation at Scale from Human-Like Contact Representations

ICRA 2026poster

Cross-embodiment dexterous grasp synthesis refers to adaptively generating and optimizing grasps for various robotic hands with different morphologies. This capability is crucial for achieving versatile robotic manipulation in diverse environments and requires substantial amounts of reliable and div…

2026

Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

CVPR 2026

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of gesture-rich data and their limited ability to infer fine-grained

Cited by 0SourceScholar
2026

Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion Models

CVPR 2026

Generating realistic human-human interactions is a challenging task that requires not only high-quality individual body and hand motions, but also coherent coordination among all interactants. Due to limitations in available data and increased learning complexity, previous methods tend to ignore han

Cited by 0SourceScholar
2026

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

ICML 2026poster

This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and free-viewpoint talking portrait synthesis, given an identity embedding or reference image, within a unified framework. Existing 2D speech-to-video diffusio…

Cited by 0SourceScholar
2025

Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidance

CVPR 2025poster

Inspired by the effectiveness of 3D Gaussian Splatting (3DGS) in reconstructing detailed 3D scenes within multi-view setups and the emergence of large 2D human foundation models, we introduce Arc2Avatar, the first SDS-based method utilizing a human face foundation model as guidance with just a singl…

Cited by 2SourcePDFScholar
2025

Deep Gaussian from Motion: Exploring 3D Geometric Foundation Models for Gaussian Splatting

NeurIPS 2025poster

Neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) are popular techniques to reconstruct and render photorealistic images. However, the prerequisite of running Structure-from-Motion (SfM) to get camera poses limits their completeness. Although previous methods can reconstruct a few unpos…

Cited by 0SourceScholar
2025

HORT: Monocular Hand-held Objects Reconstruction with Transformers

ICCV 2025poster

Reconstructing hand-held objects in 3D from monocular images remains a significant challenge in computer vision. Most existing approaches rely on implicit 3D representations, which produce overly smooth reconstructions and are time-consuming to generate explicit 3D shapes. While more recent methods…

Cited by 0SourcePDFScholar
2025

HUST: High-Fidelity Unbiased Skin Tone Estimation via Texture Quantization

ICCV 2025poster

Recent 3D facial reconstruction methods have made significant progress in shape estimation, but high-fidelity unbiased facial albedo estimation remains challenging. Existing methods rely on expensive light-stage captured data, and while they have made progress in either high-fidelity reconstruction…

2025

HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos

CVPR 2025highlight

Despite the advent in 3D hand pose estimation, current methods predominantly focus on single-image 3D hand reconstruction in the camera frame, overlooking the world-space motion of the hands. Such limitation prohibits their direct use in egocentric video settings, where hands and camera are continuo…

2025

ImHead: A Large-scale Implicit Morphable Model for Localized Head Modeling

ICCV 2025poster

Over the last years, 3D morphable models (3DMMs) have emerged as a state-of-the-art methodology for modeling and generating expressive 3D avatars. However, given their reliance on a strict topology, along with their linear nature, they struggle to represent complex full-head shapes. Following the ad…

Cited by 0SourcePDFScholar
2025

Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator

ICCV 2025poster

Sign language is a visual language that encompasses all linguistic features of natural languages and serves as the primary communication method for the deaf and hard-of-hearing communities. Although many studies have successfully adapted pretrained language models (LMs) for sign language translation…

2025

VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning

ICCV 2025poster

In recent years, video question answering based on multimodal large language models (MLLM) has garnered considerable attention, due to the benefits from the substantial advancements in LLMs. However, these models have a notable deficiency in the domains of video temporal grounding and reasoning, pos…

2025

WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild

CVPR 2025poster

In recent years, 3D hand pose estimation methods have garnered significant attention due to their extensive applications in human-computer interaction, virtual reality, and robotics. In contrast, there has been a notable gap in hand detection pipelines, posing significant challenges in constructing…

2024

AnimateMe: 4D Facial Expressions via Diffusion Models

ECCV 2024poster

"The field of photorealistic 3D avatar reconstruction and generation has garnered significant attention in recent years; however, animating such avatars remains challenging. Recent advances in diffusion models have notably enhanced the capabilities of generative models in 2D animation. In this work,…

Cited by 2SourcePDFScholar
2024

Locally Adaptive Neural 3D Morphable Models

CVPR 2024poster

We present the Locally Adaptive Morphable Model (LAMM) a highly flexible Auto-Encoder (AE) framework for learning to generate and manipulate 3D meshes. We train our architecture following a simple self-supervised training scheme in which input displacements over a set of sparse control vertices are…

2024

Neural Sign Actors: A Diffusion Model for 3D Sign Language Production from Text

CVPR 2024poster

Sign Languages (SL) serve as the primary mode of communication for the Deaf and Hard of Hearing communities. Deep learning methods for SL recognition and translation have achieved promising results. However Sign Language Production (SLP) poses a challenge as the generated motions must be realistic a…

Cited by 20SourcePDFScholar
2023

Handy: Towards a High Fidelity 3D Hand Shape and Appearance Model

CVPR 2023poster

Over the last few years, with the advent of virtual and augmented reality, an enormous amount of research has been focused on modeling, tracking and reconstructing human hands. Given their power to express human behavior, hands have been a very important, but challenging component of the human body.…

2022

Revisiting Point Cloud Simplification: A Learnable Feature Preserving Approach

ECCV 2022poster

"The recent advances in 3D sensing technology have made possible the capture of point clouds in significantly high resolution. However, increased detail usually comes at the expense of high storage, as well as computational costs in terms of processing and visualization operations. Mesh and Point Cl…

Cited by 34SourcePDFScholar
2020

Learning to Generate Customized Dynamic 3D Facial Expressions

ECCV 2020poster

Recent advances in deep learning have significantly pushed the state-of-the-art in photorealistic video animation given a single image. In this paper, we extrapolate those advances to the 3D domain, by studying 3D image-to-video translation with a particular focus on 4D facial expressions. Although…

Cited by 25SourcePDFScholar