← Search

Tae-Kyun Kim

53 accepted papers

2026

PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object Interactions

ICML 2026poster

While existing methods for reconstructing hand–object interactions have made impressive progress, they either focus on rigid or part-wise rigid objects—limiting their ability to model real-world objects (e.g., cloth, stuffed animals) that exhibit highly non-rigid deformations—or model deformable obj…

Cited by 0SourceScholar
2025

Body-Hand Modality Expertized Networks with Cross-attention for Fine-grained Skeleton Action Recognition

IROS 2025

Skeleton-based Human Action Recognition (HAR) is a vital technology in robotics and human–robot interaction. However, most existing methods concentrate primarily on full-body movements and often overlook subtle hand motions that are critical for distinguishing fine-grained actions. Recent work lever

Cited by 1SourceScholar
2025

Joint Learning of Pose Regression and Denoising Diffusion with Score Scaling Sampling for Category-level 6D Pose Estimation

ICCV 2025poster

Latest diffusion models have shown promising results in category-level 6D object pose estimation by modeling the conditional pose distribution with depth image input. The existing methods, however, suffer from slow convergence during training, learning its encoder with the diffusion denoising networ…

Cited by 0SourcePDFScholar
2025

MPMAvatar: Learning 3D Gaussian Avatars with Accurate and Robust Physics-Based Dynamics

NeurIPS 2025poster

While there has been significant progress in the field of 3D avatar creation from visual observations, modeling physically plausible dynamics of humans with loose garments remains a challenging problem. Although a few existing works address this problem by leveraging physical simulation, they suffer…

Cited by 0SourceScholar
2025

REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning

CVPR 2025poster

We present REWIND (Real-Time Egocentric Whole-Body Motion Diffusion), a one-step diffusion model for real-time, high-fidelity human motion estimation from egocentric image inputs. While an existing method for egocentric whole-body (i.e., body and hands) motion estimation is non-real-time and acausal…

Cited by 0SourcePDFScholar
2025

SRHand: Super-Resolving Hand Images and 3D Shapes via View/Pose-aware Neural Image Representations and Explicit Meshes

NeurIPS 2025poster

Reconstructing detailed hand avatars plays a crucial role in various applications. While prior works have focused on capturing high-fidelity hand geometry, they heavily rely on high-resolution multi-view image inputs and struggle to generalize on low-resolution images. Multi-view image super-resolut…

Cited by 0SourceScholar
2024

Arbitrary-Scale Image Generation and Upsampling using Latent Diffusion Model and Implicit Neural Decoder

CVPR 2024poster

Super-resolution (SR) and image generation are important tasks in computer vision and are widely adopted in real-world applications. Most existing methods however generate images only at fixed-scale magnification and suffer from over-smoothing and artifacts. Additionally they do not offer enough div…

Cited by 18SourcePDFScholar
2024

BiTT: Bi-directional Texture Reconstruction of Interacting Two Hands from a Single Image

CVPR 2024poster

Creating personalized hand avatars is important to offer a realistic experience to users on AR / VR platforms. While most prior studies focused on reconstructing 3D hand shapes some recent work has tackled the reconstruction of hand textures on top of shapes. However these methods are often limited…

2024

InterHandGen: Two-Hand Interaction Generation via Cascaded Reverse Diffusion

CVPR 2024poster

We present InterHandGen a novel framework that learns the generative prior of two-hand interaction. Sampling from our model yields plausible and diverse two-hand shapes in close interaction with or without an object. Our prior can be incorporated into any optimization or learning methods to reduce a…

Cited by 13SourcePDFScholar
2024

Multi-hypotheses Conditioned Point Cloud Diffusion for 3D Human Reconstruction from Occluded Images

NeurIPS 2024poster

3D human shape reconstruction under severe occlusion due to human-object or human-human interaction is a challenging problem. While implicit function methods capture detailed clothed shapes, they require aligned shape priors and or are weak at inpainting occluded regions given an image input. Parame…

2024

Prompt Augmentation for Self-supervised Text-guided Image Manipulation

CVPR 2024poster

Text-guided image editing finds applications in various creative and practical fields. While recent studies in image generation have advanced the field they often struggle with the dual challenges of coherent image transformation and context preservation. In response our work introduces prompt augme…

Cited by 2SourcePDFScholar
2023

3D Distillation: Improving Self-Supervised Monocular Depth Estimation on Reflective Surfaces

ICCV 2023poster

Self-supervised monocular depth estimation (SSMDE) aims at predicting the dense depth maps of monocular images, by learning to minimize a photometric loss using spatially neighboring image pairs during training. While SSMDE offers a significant scalability advantage over supervised approaches, it pe…

Cited by 11PDFScholar
2023

FourierHandFlow: Neural 4D Hand Representation Using Fourier Query Flow

NeurIPS 2023poster

Recent 4D shape representations model continuous temporal evolution of implicit shapes by (1) learning query flows without leveraging shape and articulation priors or (2) decoding shape occupancies separately for each time value. Thus, they do not effectively capture implicit correspondences between…

Cited by 5SourcePDFScholar
2023

Im2Hands: Learning Attentive Implicit Representation of Interacting Two-Hand Shapes

CVPR 2023poster

We present Implicit Two Hands (Im2Hands), the first neural implicit representation of two interacting hands. Unlike existing methods on two-hand reconstruction that rely on a parametric hand model and/or low-resolution meshes, Im2Hands can produce fine-grained geometry of two hands with high hand-to…

2023

Joint Training of Hierarchical GANs and Semantic Segmentation for Expression Translation

ICASSP 2023accepted

Manipulating images by changing only specific attributes has been a long-standing research problem. Existing methods that rely solely on a global generator often suffer from changing unwanted attributes along with the desired attributes. Although hierarchical networks consisting of global and local…

Cited by 0SourceScholar
2023

MAPConNet: Self-supervised 3D Pose Transfer with Mesh and Point Contrastive Learning

ICCV 2023poster

3D pose transfer is a challenging generation task that aims to transfer the pose of a source geometry onto a target geometry with the target identity preserved. Many prior methods require keypoint annotations to find correspondence between the source and target. Current pose transfer methods allow e…

Cited by 2PDFcodeScholar
2023

Unsupervised Contour Tracking of Live Cells by Mechanical and Cycle Consistency Losses

CVPR 2023poster

Analyzing the dynamic changes of cellular morphology is important for understanding the various functions and characteristics of live cells, including stem cells and metastatic cancer cells. To this end, we need to track all points on the highly deformable cellular contour in every frame of live cel…

2022

Pop-Out Motion: 3D-Aware Image Deformation via Learning the Shape Laplacian

CVPR 2022poster

We propose a framework that can deform an object in a 2D image as it exists in 3D space. Most existing methods for 3D-aware image manipulation are limited to (1) only changing the global scene information or depth, or (2) manipulating an object of specific categories. In this paper, we present a 3D-…

Cited by 3PDFScholar
2022

SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-Learning

NeurIPS 2022accept

Value factorisation is a useful technique for multi-agent reinforcement learning (MARL) in global reward game, however, its underlying mechanism is not yet fully understood. This paper studies a theoretical framework for value factorisation with interpretability via Shapley value theory. We generali…

2021

Adversarial Imitation Learning with Trajectorial Augmentation and Correction

ICRA 2021poster

Deep Imitation Learning requires a large number of expert demonstrations, which are not always easy to obtain, especially for complex tasks. A way to overcome this shortage of labels is through data augmentation. However, this cannot be easily applied to control tasks due to the sequential nature of…

Cited by 18SourcecodeScholar
2021

EvDistill: Asynchronous Events To End-Task Learning via Bidirectional Reconstruction-Guided Cross-Modal Knowledge Distillation

CVPR 2021poster

Event cameras sense per-pixel intensity changes and produce asynchronous event streams with high dynamic range and less motion blur, showing advantages over the conventional cameras. A hurdle of training event-based models is the lack of large qualitative labeled data. Prior works learning end-tasks…

Cited by 85PDFcodeScholar
2021

Geometry-Based Distance Decomposition for Monocular 3D Object Detection

ICCV 2021poster

Monocular 3D object detection is of great significance for autonomous driving but remains challenging. The core challenge is to predict the distance of objects in the absence of explicit depth information. Unlike regressing the distance as a single variable in most existing methods, we propose a nov…

Cited by 168PDFcodeScholar
2021

Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue System

ICLR 2021poster

Designing task-oriented dialogue systems is a challenging research topic, since it needs not only to generate utterances fulfilling user requests but also to guarantee the comprehensibility. Many previous works trained end-to-end (E2E) models with supervised learning (SL), however, the bias in annot…

2020

Active 6D Multi-Object Pose Estimation in Cluttered Scenarios with Deep Reinforcement Learning

IROS 2020

In this work, we explore how a strategic selection of camera movements can facilitate the task of 6D multi-object pose estimation in cluttered scenarios while respecting real-world constraints such as time and distance travelled, important in robotics and augmented reality applications. In the propo

Cited by 12SourceScholar
2020

Auglabel: Exploiting Word Representations to Augment Labels for Face Attribute Classification

ICASSP 2020accepted

Augmenting data in image space (eg. flipping, cropping etc) and activation space (eg. dropout) are being widely used to regularise deep neural networks and have been successfully applied on several computer vision tasks. Unlike previous works, which are mostly focused on doing augmentation in the af…

Cited by 0SourceScholar
2020

Distance-Normalized Unified Representation for Monocular 3D Object Detection

ECCV 2020poster

Monocular 3D object detection plays an important role in autonomous driving and still remains challenging. To achieve fast and accurate monocular 3D object detection, we introduce a single-stage and multi-scale framework to learn a unified representation for objects within different distance ranges,…

Cited by 65SourcePDFScholar
2020

EventSR: From Asynchronous Events to Image Reconstruction, Restoration, and Super-Resolution via End-to-End Adversarial Learning

CVPR 2020poster

Event cameras sense intensity changes and have many advantages over conventional cameras. To take advantage of event cameras, some methods have been proposed to reconstruct intensity images from event streams. However, the outputs are still in low resolution (LR), noisy, and unrealistic. The low-qua…

Cited by 123PDFcodeScholar
2020

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

ECCV 2020poster

Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction","We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art metho…

2020

Physics-Based Dexterous Manipulations with Estimated Hand Poses and Residual Reinforcement Learning

IROS 2020poster

Dexterous manipulation of objects in virtual environments with our bare hands, by using only a depth sensor and a state-of-the-art 3D hand pose estimator (HPE), is challenging. While virtual environments are ruled by physics, e.g. object weights and surface frictions, the absence of force feedback m…

Cited by 62SourceScholar
2020

Weakly-Supervised Domain Adaptation via GAN and Mesh Model for Estimating 3D Hand Poses Interacting Objects

CVPR 2020oral

Despite recent successes in hand pose estimation, there yet remain challenges on RGB-based 3D hand pose estimation (HPE) under hand-object interaction (HOI) scenarios where severe occlusions and cluttered backgrounds exhibit. Recent RGB HOI benchmarks have been collected either in real or synthetic…

Cited by 101PDFcodeScholar
2019

Pushing the Envelope for RGB-Based Dense 3D Hand Pose Estimation via Neural Rendering

CVPR 2019poster

Estimating 3D hand meshes from single RGB images is challenging, due to intrinsic 2D-3D mapping ambiguities and limited training data. We adopt a compact parametric 3D hand model that represents deformable and articulated hand meshes. To achieve the model fitting to RGB images, we investigate and co…

Cited by 268PDFScholar
2018

Augmented Skeleton Space Transfer for Depth-Based Hand Pose Estimation

CVPR 2018poster

Crucial to the success of training a depth-based 3D hand pose estimator (HPE) is the availability of comprehensive datasets covering diverse camera perspectives, shapes, and pose variations. However, collecting such annotated datasets is challenging. We propose to complete existing databases by gene…

Cited by 101SourcePDFScholar
2018

BOP: Benchmark for 6D Object Pose Estimation

ECCV 2018poster

We propose a benchmark for 6D pose estimation of a rigid object from a single RGB-D input image. The training data consists of a texture-mapped 3D object model or images of the object in known 6D poses. The benchmark comprises of: i) eight datasets in a unified format that cover different practical…

2018

Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals

CVPR 2018poster

In this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods o…

Cited by 277SourcePDFScholar
2018

First-Person Hand Action Benchmark With RGB-D Videos and 3D Hand Pose Annotations

CVPR 2018poster

In this work we study the use of 3D hand poses to recognize first-person dynamic hand actions interacting with 3D objects. Towards this goal, we collected RGB-D video sequences comprised of more than 100K frames of 45 daily hand action categories, involving 26 different objects in several hand conf…

Cited by 672SourcePDFScholar
2018

Semi-supervised Adversarial Learning to Generate Photorealistic Face Images of New Identities from 3D Morphable Model

ECCV 2018poster

We propose a novel end-to-end semi-supervised adversarial framework to generate photorealistic face images of new identities with a wide range of expressions, poses, and illuminations conditioned by synthetic images sampled from a 3D morphable model. Previous adversarial style-transfer methods eithe…

2017

BigHand2.2M Benchmark: Hand Pose Dataset and State of the Art Analysis

CVPR 2017poster

In this paper we introduce a large-scale hand pose dataset, collected using a novel capture method. Existing datasets are either generated synthetically or captured using depth sensors: synthetic datasets exhibit a certain level of appearance difference from real depth images, and real datasets are…

Cited by 329PDFScholar
2017

Learning and Refining of Privileged Information-Based RNNs for Action Recognition From Depth Sequences

CVPR 2017poster

Existing RNN-based approaches for action recognition from depth sequences require either skeleton joints or hand-crafted depth features as inputs. An end-to-end manner, mapping from raw depth maps to action classes, is non-trivial to design due to the fact that: 1) single channel map lacks texture t…

Cited by 102PDFScholar
2017

Pose Guided RGBD Feature Learning for 3D Object Pose Estimation

ICCV 2017poster

In this paper we examine the effects of using object poses as guidance to learning robust features for 3D object pose estimation. Previous works have focused on learning feature embeddings based on metric learning with triplet comparisons and rely only on the qualitative distinction of similar and d…

Cited by 79PDFScholar
2017

Transition Forests: Learning Discriminative Temporal Transitions for Action Recognition and Detection

CVPR 2017poster

A human action can be seen as transitions between one's body poses over time, where the transition depicts a temporal relation between two poses. Recognizing actions thus involves learning a classifier sensitive to these pose transitions as well as to static poses. In this paper, we introduce a nove…

Cited by 102PDFScholar
2016

Iterative Hough Forest with Histogram of Control Points for 6 DoF object registration from depth images

IROS 2016poster

State-of-the-art techniques proposed for 6D object pose recovery depend on occlusion-free point clouds to accurately register objects in 3D space. To reduce this dependency, we introduce a novel architecture called Iterative Hough Forest with Histogram of Control Points that is capable of estimating…

Cited by 16SourceScholar
2016

Recovering 6D Object Pose and Predicting Next-Best-View in the Crowd

CVPR 2016poster

Object detection and 6D pose estimation in the crowd (scenes with multiple object instances, severe foreground occlusions and background distractors), has become an important problem in many rapidly evolving technological areas such as robotics and augmented reality. Single shot-based 6D pose estima…

Cited by 275PDFScholar
2015

Conditional Convolutional Neural Network for Modality-Aware Face Recognition

ICCV 2015poster

Faces in the wild are usually captured with various poses, illuminations and occlusions, and thus inherently multimodally distributed in many tasks. We propose a conditional Convolutional Neural Network, named as c-CNN, to handle multimodal face recognition. Different from traditional CNN that adopt…

Cited by 111PDFScholar
2015

Opening the Black Box: Hierarchical Sampling Optimization for Estimating Human Hand Pose

ICCV 2015oral

We address the problem of hand pose estimation, formulated as an inverse problem. Typical approaches optimize an energy function over pose parameters using a `black box' image generation procedure. This procedure knows little about either the relationships between the parameters or the form of the…

Cited by 170PDFScholar