← Search

Sanghoon Lee

22 accepted papers

2026

DipGuava: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular Video

AAAI 2026technical

While recent 3D head avatar creation methods attempt to animate facial dynamics, they often fail to capture personalized details, limiting realism and expressiveness. To fill this gap, we present DipGuava (Disentangled and Personalized Gaussian UV Avatar), a novel 3D Gaussian head avatar creation me

Cited by 0SourcePDFScholar
2025

MBTI: Masked Blending Transformers with Implicit Positional Encoding for Frame-rate Agnostic Motion Estimation

ICCV 2025poster

Human motion estimation models typically assume a fixed number of input frames, making them sensitive to variations in frame rate and leading to inconsistent motion predictions across different temporal resolutions. This limitation arises because input frame rates inherently determine the temporal g…

Cited by 0SourcePDFScholar
2025

SDAS: Semantic Data Acquisition System for Minimizing Redundancy and Maximizing Diversity

AAAI 2025technical

In this paper, we propose SDAS, a new motion data assessment and storage system designed to acquire new motion data with reduced redundancy and maximizing diversity. SDAS collects data in the field, retrieves the most similar data from the database in real-time, and provides visualization tools that…

Cited by 0SourcePDFScholar
2025

Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURA

ICCV 2025poster

Amodal segmentation aims to infer the complete shape of occluded objects, even when the occluded region's appearance is unavailable. However, current amodal segmentation methods lack the capability to interact with users through text input and struggle to understand or reason about implicit and comp…

2024

AVIN-Chat: An Audio-Visual Interactive Chatbot System with Emotional State Tuning

IJCAI 2024poster

This work presents an audio-visual interactive chatbot (AVIN-Chat) system that allows users to have face-to-face conversations with 3D avatars in real-time. Compared to the previous chatbot services, which provide text-only or speech-only communications, the proposed AVIN-Chat can offer audio-visual…

2024

InViTe: Individual Virtual Transfer for Personalized 3D Face Generation System

IJCAI 2024poster

With the expansion of the virtual communication industry using VR/AR, it has attracted increasing attention to enable users to represent their personalities in a 3D avatar. As the face of 3D avatars plays a crucial role in conveying human personality, a system that generates and manipulates 3D faces…

Cited by 0SourcePDFScholar
2024

Speech-Driven Emotional 3d Talking Face Animation Using Emotional Embeddings

ICASSP 2024accepted

Existing emotional talking 3D facial animation primarily focus on animating emotional faces using a specific emotion condition. However, in real-world situations, no one consistently speaks with just one emotion. Thus, previous emotion-based approaches have very limited applicability in real-world a…

Cited by 0SourceScholar
2024

TurboHopp: Accelerated Molecule Scaffold Hopping with Consistency Models

NeurIPS 2024poster

Navigating the vast chemical space of druggable compounds is a formidable challenge in drug discovery, where generative models are increasingly employed to identify viable candidates. Conditional 3D structure-based drug design (3D-SBDD) models, which take into account complex three-dimensional inter…

2023

Camera-Driven Representation Learning for Unsupervised Domain Adaptive Person Re-identification

ICCV 2023poster

We present a novel unsupervised domain adaption method for person re-identification (reID) that generalizes a model trained on a labeled source domain to an unlabeled target domain. We introduce a camera-driven curriculum learning (CaCL) framework that leverages camera labels of person images to tra…

Cited by 40PDFScholar
2022

A Brand New Dance Partner: Music-Conditioned Pluralistic Dancing Controlled by Multiple Dance Genres

CVPR 2022poster

When coming up with phrases of movement, choreographers all have their habits as they are used to their skilled dance genres. Therefore, they tend to return certain patterns of the dance genres that they are familiar with. What if artificial intelligence could be used to help choreographers blend da…

Cited by 51PDFcodeScholar
2022

Decomposed Knowledge Distillation for Class-Incremental Semantic Segmentation

NeurIPS 2022accept

Class-incremental semantic segmentation (CISS) labels each pixel of an image with a corresponding object/stuff class continually. To this end, it is crucial to learn novel classes incrementally without forgetting previously learned knowledge. Current CISS methods typically use a knowledge distillati…

Cited by 38SourcePDFScholar
2022

OIMNet++: Prototypical Normalization and Localization-Aware Learning for Person Search

ECCV 2022poster

"We address the task of person search, that is, localizing and re-identifying query persons from a set of raw scene images. Recent approaches are typically built upon OIMNet, a pioneer work on person search, that learns joint person representations for performing both detection and person re-identif…

2021

HVPR: Hybrid Voxel-Point Representation for Single-Stage 3D Object Detection

CVPR 2021poster

We address the problem of 3D object detection, that is, estimating 3D object bounding boxes from point clouds. 3D object detection methods exploit either voxel-based or point-based features to represent 3D objects in a scene. Voxel-based features are efficient to extract, while they fail to preserve…

Cited by 169PDFcodeScholar
2021

Learning by Aligning: Visible-Infrared Person Re-Identification Using Cross-Modal Correspondences

ICCV 2021poster

We address the problem of visible-infrared person re-identification (VI-reID), that is, retrieving a set of person images, captured by visible or infrared cameras, in a cross-modal setting. Two main challenges in VI-reID are intra-class variations across person images, and cross-modal discrepancies…

Cited by 246PDFScholar
2019

A Deep Cybersickness Predictor Based on Brain Signal Analysis for Virtual Reality Contents

ICCV 2019poster

What if we could interpret the cognitive state of a user while experiencing a virtual reality (VR) and estimate the cognitive state from a visual stimulus? In this paper, we address the above question by developing an electroencephalography (EEG) driven VR cybersickness prediction model. The EEG dat…

Cited by 101PDFScholar
2019

GraphX-Convolution for Point Cloud Deformation in 2D-to-3D Conversion

ICCV 2019poster

In this paper, we present a novel deep method to reconstruct a point cloud of an object from a single still image. Prior arts in the field struggle to reconstruct an accurate and scalable 3D model due to either the inefficient and expensive 3D representations, the dependency between the output and n…

Cited by 42PDFcodeScholar
2018

Deep Video Quality Assessor: From Spatio-temporal Visual Sensitivity to A Convolutional Neural Aggregation Network

ECCV 2018poster

Incorporating spatio-temporal human visual perception into video quality assessment (VQA) remains a formidable issue. Previous statistical or computational models of spatio-temporal perception have limitations to be applied to the general VQA algorithms. In this paper, we propose a novel full-refere…

Cited by 151SourcePDFScholar
2017

Ensemble Deep Learning for Skeleton-Based Action Recognition Using Temporal Sliding LSTM Networks

ICCV 2017poster

This paper addresses the problems of feature representation of skeleton joints and the modeling of temporal dynamics to recognize human actions. Traditional methods generally use relative coordinate systems dependent on some joints, and model only the long-term dependency, while excluding short-term…

Cited by 506PDFScholar