← Search

Seungryul Baek

20 accepted papers

2026

HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models

CVPR 2026

Understanding the fine-grained articulation of human hands is critical in high-stakes settings such as robot-assisted surgery, chip manufacturing, and AR/VR-based human-AI interaction. Despite achieving near-human performance on general vision-language benchmarks, current vision-language models (VLM

Cited by 0SourceScholar
2025

BIGS: Bimanual Category-agnostic Interaction Reconstruction from Monocular Videos via 3D Gaussian Splatting

CVPR 2025poster

Reconstructing 3Ds of hand-object interaction (HOI) is a fundamental problem that can find numerous applications. Despite recent advances, there is no comprehensive pipeline yet for bimanual class-agnostic interaction reconstruction from a monocular RGB video, where two hands and an unknown object a…

Cited by 0SourcePDFScholar
2025

Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection

ICCV 2025poster

We introduce a deepfake video detection approach that exploits pixel-wise temporal inconsistencies, which traditional spatial frequency-based detectors often overlook. The traditional detectors represent temporal information merely by stacking spatial frequency spectra across frames, resulting in th…

2025

PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation

CVPR 2025poster

We study multi-dataset training (MDT) for pose estimation, where skeletal heterogeneity presents a unique challenge that existing methods have yet to address. In traditional domains, e.g. regression and classification, MDT typically relies on dataset merging or multi-head supervision. However, the d…

2025

QORT-Former: Query-optimized Real-time Transformer for Understanding Two Hands Manipulating Objects

AAAI 2025technical

Significant advancements have been achieved in the realm of understanding poses and interactions of two hands manipulating an object. The emergence of augmented reality (AR) and virtual reality (VR) technologies has heightened the demand for real-time performance in these applications. However, curr…

Cited by 1SourcePDFScholar
2025

Text2Relight: Creative Portrait Relighting with Text Guidance

AAAI 2025technical

We present a lighting-aware image editing pipeline that, given a portrait image and a text prompt, performs single image relighting. Our model modifies the lighting and color of both the foreground and background to align with the provided text description. The unbounded nature in creativeness of a…

Cited by 1SourcePDFScholar
2024

Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects

ECCV 2024poster

"We interact with the world with our hands and see it through our own (egocentric) perspective. A holistic understanding of such interactions from egocentric views is important for tasks in robotics, AR/VR, action recognition and motion generation. Accurately reconstructing such interactions in is c…

2024

Class-Wise Buffer Management for Incremental Object Detection: An Effective Buffer Training Strategy

ICASSP 2024accepted

Class incremental learning aims to solve a problem that arises when continuously adding unseen class instances to an existing model This approach has been extensively studied in the context of image classification; however its applicability to object detection is not well established yet. Existing f…

Cited by 0SourceScholar
2024

Exploiting Style Latent Flows for Generalizing Deepfake Video Detection

CVPR 2024poster

This paper presents a new approach for the detection of fake videos based on the analysis of style latent vectors and their abnormal behavior in temporal changes in the generated videos. We discovered that the generated facial videos suffer from the temporal distinctiveness in the temporal changes o…

Cited by 34SourcePDFScholar
2024

SDDGR: Stable Diffusion-based Deep Generative Replay for Class Incremental Object Detection

CVPR 2024highlight

In the field of class incremental learning (CIL) generative replay has become increasingly prominent as a method to mitigate the catastrophic forgetting alongside the continuous improvements in generative models. However its application in class incremental object detection (CIOD) has been significa…

Cited by 28SourcePDFScholar
2024

Text2HOI: Text-guided 3D Motion Generation for Hand-Object Interaction

CVPR 2024poster

This paper introduces the first text-guided work for generating the sequence of hand-object interaction in 3D. The main challenge arises from the lack of labeled data where existing ground-truth datasets are nowhere near generalizable in interaction type and object category which inhibits the modeli…

2023

Transformer-Based Unified Recognition of Two Hands Manipulating Objects

CVPR 2023poster

Understanding the hand-object interactions from an egocentric video has received a great attention recently. So far, most approaches are based on the convolutional neural network (CNN) features combined with the temporal encoding via the long short-term memory (LSTM) or graph convolution network (GC…

2022

Multi-Person 3D Pose and Shape Estimation via Inverse Kinematics and Refinement

ECCV 2022poster

"Estimating 3D poses and shapes in the form of meshes from monocular RGB images is challenging. Obviously, it is more difficult than estimating 3D poses only in the form of skeletons or heatmaps. When interacting persons are involved, the 3D mesh reconstruction becomes more challenging due to the am…

2020

Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction

ECCV 2020poster

Articulations, Shapes and Objects for 3D Hand Pose Estimation under Hand-Object Interaction","We study how well different types of approaches generalise in the task of 3D hand pose estimation under single hand scenarios and hand-object interaction. We show that the accuracy of state-of-the-art metho…

2020

Weakly-Supervised Domain Adaptation via GAN and Mesh Model for Estimating 3D Hand Poses Interacting Objects

CVPR 2020oral

Despite recent successes in hand pose estimation, there yet remain challenges on RGB-based 3D hand pose estimation (HPE) under hand-object interaction (HOI) scenarios where severe occlusions and cluttered backgrounds exhibit. Recent RGB HOI benchmarks have been collected either in real or synthetic…

Cited by 101PDFcodeScholar
2019

Pushing the Envelope for RGB-Based Dense 3D Hand Pose Estimation via Neural Rendering

CVPR 2019poster

Estimating 3D hand meshes from single RGB images is challenging, due to intrinsic 2D-3D mapping ambiguities and limited training data. We adopt a compact parametric 3D hand model that represents deformable and articulated hand meshes. To achieve the model fitting to RGB images, we investigate and co…

Cited by 268PDFScholar
2018

Augmented Skeleton Space Transfer for Depth-Based Hand Pose Estimation

CVPR 2018poster

Crucial to the success of training a depth-based 3D hand pose estimator (HPE) is the availability of comprehensive datasets covering diverse camera perspectives, shapes, and pose variations. However, collecting such annotated datasets is challenging. We propose to complete existing databases by gene…

Cited by 101SourcePDFScholar
2018

First-Person Hand Action Benchmark With RGB-D Videos and 3D Hand Pose Annotations

CVPR 2018poster

In this work we study the use of 3D hand poses to recognize first-person dynamic hand actions interacting with 3D objects. Towards this goal, we collected RGB-D video sequences comprised of more than 100K frames of 45 daily hand action categories, involving 26 different objects in several hand conf…

Cited by 672SourcePDFScholar