← Search

Hyung Jin Chang

53 accepted papers

2026

AISPO: Enhancing Depth Reliability for Robotic Manipulation of Non-Lambertian Objects via Affine-Invariant Shape Prior

RA-L 2026

Reliable depth perception is critical for robotic manipulation, especially for non-Lambertian objects such as transparent or highly specular surfaces, where raw depth measurements are often corrupted or missing. These failures frequently propagate to motion planning, resulting in invalid grasp poses

Cited by 0SourceScholar
2026

FlatLab: A Unified Methodology Framework and Simulation-Based Benchmark for Robotic Manipulation of Flat Objects

ICML 2026poster

Robotic manipulation of flat objects is challenging due to the ungraspable configurations and strong variations in object geometry and material. Existing methods rely on heuristic pre-manipulation and are often evaluated in closed settings with limited generalization. We propose a unified framework …

Cited by 0SourcecodeScholar
2026

Force-Aware 3D Contact Modeling for Stable Grasp Generation

AAAI 2026technical

Contact-based grasp generation plays a crucial role in various applications. Recent methods typically focus on the geometric structure of objects, producing grasps with diverse hand poses and plausible contact points. However, these approaches often overlook the physical attributes of the grasp, spe

Cited by 0SourcePDFScholar
2026

GraspALL: Adaptive Structural Compensation from Illumination Variation for Robotic Garment Grasping in Any Low-Light Conditions

CVPR 2026

Achieving accurate garment grasping under dynamically changing illumination is crucial for all-day operation of service robots. However, the reduced illumination in low-light scenes severely degrades garment structural features, leading to a significant drop in grasping robustness. Existing methods

Cited by 0SourcecodeScholar
2026

RTGaze: Real-Time 3D-Aware Gaze Redirection from a Single Image

AAAI 2026technical

Gaze redirection methods aim to generate realistic human face images with controllable eye movement. However, recent methods often struggle with 3D consistency, efficiency, or quality, limiting their practical applications. In this work, we propose RTGaze, a real-time and high-quality gaze redirecti

Cited by 0SourcePDFScholar
2025

3D Prior Is All You Need: Cross-Task Few-shot 2D Gaze Estimation

CVPR 2025poster

3D and 2D gaze estimation share the fundamental objective of capturing eye movements but are traditionally treated as two distinct research domains. In this paper, we introduce a novel cross-task few-shot 2D gaze estimation approach, aiming to adapt a pre-trained 3D gaze estimation network for 2D ga…

Cited by 0SourcePDFScholar
2025

AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and Segmentation

ICCV 2025poster

The challenge of multimodal semantic segmentation lies in establishing semantically consistent and segmentable multimodal fusion features under conditions of significant visual feature discrepancies. Existing methods commonly construct cross-modal self-attention fusion frameworks or introduce additi…

2025

Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics

AAAI 2025technical

With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise seen actions on unseen objects due to the limitations in re…

Cited by 0SourcePDFScholar
2025

DarkSeg: Infrared-Driven Semantic Segmentation for Garment Grasping Detection in Low-Light Conditions

IROS 2025

Garment grasping in low-light environments is a critical challenge for domestic intelligent robots, yet existing research has not sufficiently addressed this issue. In low-light conditions, the scarcity of visual features due to insufficient illumination causes different categories of garments to ex

Cited by 1SourcecodeScholar
2025

High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation

ICCV 2025poster

Modeling high-resolution spatiotemporal representations, including both global dynamic contexts (e.g., holistic human motion tendencies) and local motion details (e.g., high-frequency changes of keypoints), is essential for video-based human pose estimation (VHPE). Current state-of-the-art methods t…

Cited by 0SourcePDFScholar
2025

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models

ICCV 2025poster

Continual learning enables pre-trained generative vision-language models (VLMs) to incorporate knowledge from new tasks without retraining data from previous ones. Recent methods update a visual projector to translate visual information for new tasks, connecting pre-trained vision encoders with larg…

Cited by 0SourcePDFScholar
2025

PersonaBooth: Personalized Text-to-Motion Generation

CVPR 2025poster

This paper introduces Motion Personalization, a new task that generates personalized motions aligned with text descriptions using several basic motions containing Persona. To support this novel task, we introduce a new large-scale motion dataset called PerMo (PersonaMotion), which captures the uniqu…

Cited by 1SourcePDFScholar
2025

PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation

CVPR 2025poster

We study multi-dataset training (MDT) for pose estimation, where skeletal heterogeneity presents a unique challenge that existing methods have yet to address. In traditional domains, e.g. regression and classification, MDT typically relies on dataset merging or multi-head supervision. However, the d…

2025

Single-view Image to Novel-view Generation for Hand-Object Interactions

AAAI 2025technical

Hand-object interaction modeling from a single RGB image is a significantly challenging task. Previous works typically reconstruct hand-object interactions as texture-less meshes, ignoring photo-realistic image generation. In this work, we introduce the HO123, a novel method to synthesize novel-view…

Cited by 0SourcePDFScholar
2024

Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects

ECCV 2024poster

"We interact with the world with our hands and see it through our own (egocentric) perspective. A holistic understanding of such interactions from egocentric views is important for tasks in robotics, AR/VR, action recognition and motion generation. Accurately reconstructing such interactions in is c…

2024

GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose Refinement

CVPR 2024poster

Object pose refinement is essential for robust object pose estimation. Previous work has made significant progress towards instance-level object pose refinement. Yet category-level pose refinement is a more challenging problem due to large shape variations within a category and the discrepancies bet…

Cited by 4SourcePDFScholar
2024

MoST: Motion Style Transformer Between Diverse Action Contents

CVPR 2024poster

While existing motion style transfer methods are effective between two motions with identical content their performance significantly diminishes when transferring style between motions with different contents. This challenge lies in the lack of clear separation between content and style of a motion.…

2024

NL2Contact: Natural Language Guided 3D Hand-Object Contact Modeling with Diffusion Model

ECCV 2024oral

"Modeling the physical contacts between the hand and object is standard for refining inaccurate hand poses and generating novel human grasp in 3D hand-object reconstruction. However, existing methods rely on geometric constraints that cannot be specified or controlled. This paper introduces a novel…

Cited by 2SourcePDFScholar
2024

What Do You See in Vehicle? Comprehensive Vision Solution for In-Vehicle Gaze Estimation

CVPR 2024poster

Driver's eye gaze holds a wealth of cognitive and intentional cues crucial for intelligent vehicles. Despite its significance research on in-vehicle gaze estimation remains limited due to the scarcity of comprehensive and well-annotated datasets in real driving scenarios. In this paper we present th…

Cited by 10SourcePDFScholar
2023

Clothes Grasping and Unfolding Based on RGB-D Semantic Segmentation

ICRA 2023poster

Clothes grasping and unfolding is a core step in robotic-assisted dressing. Most existing works leverage depth images of clothes to train a deep learning-based model to recognize suitable grasping points. These methods often utilize physics engines to synthesize depth images to reduce the cost of re…

Cited by 5SourceScholar
2023

DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose Estimation

ICCV 2023poster

Denoising diffusion probabilistic models that were initially proposed for realistic image generation have recently shown success in various perception tasks (e.g., object detection and image segmentation) and are increasingly gaining attention in computer vision. However, extending such models to mu…

Cited by 43PDFScholar
2023

GazeNeRF: 3D-Aware Gaze Redirection With Neural Radiance Fields

CVPR 2023poster

We propose GazeNeRF, a 3D-aware method for the task of gaze redirection. Existing gaze redirection methods operate on 2D images and struggle to generate 3D consistent results. Instead, we build on the intuition that the face region and eye balls are separate 3D structures that move in a coordinated…

2023

HS-Pose: Hybrid Scope Feature Extraction for Category-Level Object Pose Estimation

CVPR 2023poster

In this paper, we focus on the problem of category-level object pose estimation, which is challenging due to the large intra-category shape variation. 3D graph convolution (3D-GC) based methods have been widely used to extract local geometric features, but they have limitations for complex shaped ob…

2023

Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in Video

CVPR 2023poster

Temporal modeling is crucial for multi-frame human pose estimation. Most existing methods directly employ optical flow or deformable convolution to predict full-spectrum motion fields, which might incur numerous irrelevant cues, such as a nearby person or background. Without further efforts to excav…

Cited by 24SourcePDFScholar
2023

Spectral Graphormer: Spectral Graph-Based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color Images

ICCV 2023poster

We propose a novel transformer-based framework that reconstructs two high fidelity hands from multi-view RGB images. Unlike existing hand pose estimation methods, where one typically trains a deep network to regress hand model parameters from single RGB image, we consider a more challenging problem…

Cited by 4PDFScholar
2023

Switching Temporary Teachers for Semi-Supervised Semantic Segmentation

NeurIPS 2023poster

The teacher-student framework, prevalent in semi-supervised semantic segmentation, mainly employs the exponential moving average (EMA) to update a single teacher's weights based on the student's. However, EMA updates raise a problem in that the weights of the teacher and student are getting coupled,…

2022

Collaborative Learning for Hand and Object Reconstruction With Attention-Guided Graph Convolution

CVPR 2022poster

Estimating the pose and shape of hands and objects under interaction finds numerous applications including augmented and virtual reality. Existing approaches for hand and object reconstruction require explicitly defined physical constraints and known objects, which limits its application domains. Ou…

Cited by 46PDFScholar
2022

Contrastive Vicinal Space for Unsupervised Domain Adaptation

ECCV 2022poster

"Recent unsupervised domain adaptation methods have utilized vicinal space between the source and target domains. However, the equilibrium collapse of labels, a problem where the source labels are dominant over the target labels in the predictions of vicinal instances, has never been addressed. In t…

2022

Global-Local Motion Transformer for Unsupervised Skeleton-Based Action Learning

ECCV 2022poster

"We propose a new transformer model for the task of unsupervised learning of skeleton motion sequences. The existing transformer model utilized for unsupervised skeleton-based action learning is learned the instantaneous velocity of each joint from adjacent frames without global motion information.…

2022

S2Contact: Graph-Based Network for 3D Hand-Object Contact Estimation with Semi-Supervised Learning

ECCV 2022poster

"Being able to reason about the physical contacts between hands and objects is crucial in understanding hand-object manipulation. However, despite the efforts in accurate 3D annotations in hand and object datasets, there still exist gaps in 3D hand and object reconstructions. Recent works leverage c…

Cited by 21SourcePDFScholar
2022

TP-AE: Temporally Primed 6D Object Pose Tracking with Auto-Encoders

ICRA 2022poster

Fast and accurate tracking of an object's motion is one of the key functionalities of a robotic system for achieving reliable interaction with the environment. This paper focuses on the instance-level six-dimensional (6D) pose tracking problem with a symmetric and textureless object under occlusion.…

Cited by 9SourceScholar
2022

Towards Generic 3D Tracking in RGBD Videos: Benchmark and Baseline

ECCV 2022poster

"Tracking in 3D scenes is gaining momentum because of its numerous applications in robotics, autonomous driving, and scene understanding. Currently, 3D tracking is limited to specific model-based approaches involving point clouds, which impedes 3D trackers from applying in natural 3D scenes. RGBD se…

2021

Apparently Irrational Choice as Optimal Sequential Decision Making

AAAI 2021technical

In this paper, we propose a normative approach to modeling apparently human irrational decision making (cognitive biases) that makes use of inherently rational computational mechanisms. We view preferential choice tasks as sequential decision making problems and formulate them as Partially Observabl…

2021

Class-Attentive Diffusion Network for Semi-Supervised Classification

AAAI 2021technical

Recently, graph neural networks for semi-supervised classification have been widely studied. However, existing methods only use the information of limited neighbors and do not deal with the inter-class connections in graphs. In this paper, we propose Adaptive aggregation with Class-Attentive Diffusi…

2021

FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation Mechanism

CVPR 2021poster

In this paper, we focus on category-level 6D pose and size estimation from a monocular RGB-D image. Previous methods suffer from inefficient category-level pose feature extraction, which leads to low accuracy and inference speed. To tackle this problem, we propose a fast shape-based network (FS-Net)…

Cited by 197PDFcodeScholar
2021

FixBi: Bridging Domain Spaces for Unsupervised Domain Adaptation

CVPR 2021poster

Unsupervised domain adaptation (UDA) methods for learning domain invariant representations have achieved remarkable progress. However, most of the studies were based on direct adaptation from the source domain to the target domain and have suffered from large domain discrepancies. In this paper, we…

Cited by 290PDFcodeScholar
2021

Unsupervised Hyperbolic Representation Learning via Message Passing Auto-Encoders

CVPR 2021poster

Most of the existing literature regarding hyperbolic embedding concentrate upon supervised learning, whereas the use of unsupervised hyperbolic embedding is less well explored. In this paper, we analyze how unsupervised tasks can benefit from learned representations in hyperbolic space. To explore h…

Cited by 39PDFcodeScholar
2021

VaB-AL: Incorporating Class Imbalance and Difficulty With Variational Bayes for Active Learning

CVPR 2021poster

Active Learning for discriminative models has largely been studied with the focus on individual samples, with less emphasis on how classes are distributed or which classes are hard to deal with. In this work, we show that this is harmful. We propose a method based on the Bayes' rule, that can natura…

Cited by 57PDFScholar
2020

G2L-Net: Global to Local Network for Real-Time 6D Pose Estimation With Embedding Vector Features

CVPR 2020poster

In this paper, we propose a novel real-time 6D object pose estimation framework, named G2L-Net. Our network operates on point clouds from RGB-D detection in a divide-and-conquer fashion. Specifically, our network consists of three steps. First, we extract the coarse object point cloud from the RGB-D…

Cited by 133PDFcodeScholar
2020

SeqHAND: RGB-Sequence-Based 3D Hand Pose and Shape Estimation

ECCV 2020poster

3D hand pose estimation based on RGB images has been studied for a long time. Most of the studies, however, have performed frame-by-frame estimation based on independent static images. In this paper, we attempt to not only consider the appearance of a hand but incorporate the temporal movement infor…

Cited by 67SourcePDFScholar
2019

Symmetric Graph Convolutional Autoencoder for Unsupervised Graph Representation Learning

ICCV 2019poster

We propose a symmetric graph convolutional autoencoder which produces a low-dimensional latent representation from a graph. In contrast to the existing graph autoencoders with asymmetric decoder parts, the proposed autoencoder has a newly designed decoder which builds a completely symmetric autoenco…

Cited by 319PDFScholar
2018

Context-Aware Deep Feature Compression for High-Speed Visual Tracking

CVPR 2018poster

We propose a new context-aware correlation filter based tracking framework to achieve both high computational speed and state-of-the-art performance among real-time trackers. The major contribution to the high computational speed lies in the proposed deep feature compression that is achieved by a co…

2018

RT-GENE: Real-Time Eye Gaze Estimation in Natural Environments

ECCV 2018poster

In this work, we consider the problem of robust gaze estimation in natural environments. Large camera-to-subject distances and high variations in head pose and eye gaze angles are common in such environments. This leads to two main shortfalls in state-of-the-art methods for gaze estimation: hindered…

Cited by 423SourcePDFScholar
2018

Transferring Visuomotor Learning from Simulation to the Real World for Robotics Manipulation Tasks

IROS 2018poster

Hand-eye coordination is a requirement for many manipulation tasks including grasping and reaching. However, accurate hand-eye coordination has shown to be especially difficult to achieve in complex robots like the iCub humanoid. In this work, we solve the hand-eye coordination task using a visuomot…

Cited by 19SourceScholar
2017

Attentional Correlation Filter Network for Adaptive Visual Tracking

CVPR 2017poster

We propose a new tracking framework with an attentional mechanism that chooses a subset of the associated correlation filters for increased robustness and computational efficiency. The subset of filters is adaptively selected by a deep attentional network according to the dynamic properties of the t…

Cited by 388PDFScholar
2017

Variational Autoencoded Regression: High Dimensional Regression of Visual Data on Complex Manifold

CVPR 2017poster

This paper proposes a new high dimensional regression method by merging Gaussian process regression into a variational autoencoder framework. In contrast to other regression methods, the proposed method focuses on the case where output responses are on a complex high dimensional manifold, such as im…

Cited by 36PDFScholar
2016

Iterative path optimisation for personalised dressing assistance using vision and force information

IROS 2016poster

We propose an online iterative path optimisation method to enable a Baxter humanoid robot to assist human users to dress. The robot searches for the optimal personalised dressing path using vision and force sensor information: vision information is used to recognise the human pose and model the move…

Cited by 82SourceScholar
2016

Kinematic Structure Correspondences via Hypergraph Matching

CVPR 2016poster

In this paper, we present a novel framework for finding the kinematic structure correspondence between two objects in videos via hypergraph matching. In contrast to prior appearance and graph alignment based matching methods which have been applied among two similar static images, the proposed metho…

Cited by 17PDFScholar
2016

Visual Tracking Using Attention-Modulated Disintegration and Integration

CVPR 2016poster

In this paper, we present a novel attention-modulated visual tracking algorithm that decomposes an object into multiple cognitive units, and trains multiple elementary trackers in order to modulate the distribution of attention according to various feature and kernel types. In the integration stage…

Cited by 215PDFScholar
2015

Unsupervised Learning of Complex Articulated Kinematic Structures Combining Motion and Skeleton Information

CVPR 2015poster

In this paper we present a novel framework for unsupervised kinematic structure learning of complex articulated objects from a single-view image sequence. In contrast to prior motion information based methods, which estimate relatively simple articulations, our method can generate arbitrarily comple…

Cited by 22SourcePDFScholar