← Search

Taeyoung Kim

10 accepted papers

2026

Contrastive Representation Regularization for Vision-Language-Action Models

ICML 2026poster

Vision-Language-Action (VLA) models have shown strong capabilities in robot manipulation by leveraging rich representations from pre-trained Vision-Language Models (VLMs). However, their representations arguably remain suboptimal, lacking sensitivity to robotic signals such as control actions and pr…

Cited by 0SourceScholar
2026

Dynamics-Aware Planning Representation for Zero-Shot Reinforcement Learning (Student Abstract)

AAAI 2026technical

Offline Zero-Shot Reinforcement Learning requires an agent to solve unseen tasks using only a fixed offline dataset without explicit rewards. A central challenge is learning representations that capture both high-level long-term planning and low-level physical dynamics. We propose a novel framework,

Cited by 0SourcePDFScholar
2026

HAMLET: Switch Your Vision-Language-Action Model into a History-Aware Policy

ICLR 2026poster

Inherently, robotic manipulation tasks are history-dependent: leveraging past context could be beneficial. However, most existing Vision-Language-Action models (VLAs) have been designed without considering this aspect, i.e., they rely solely on the current observation, ignoring preceding context. In…

Cited by 0SourceScholar
2024

Cluster-Based Sampling in Hindsight Experience Replay for Robotic Tasks (Student Abstract)

AAAI 2024technical

In multi-goal reinforcement learning with a sparse binary reward, training agents is particularly challenging, due to a lack of successful experiences. To solve this problem, hindsight experience replay (HER) generates successful experiences even from unsuccessful ones. However, generating successfu…

Cited by 0SourcePDFScholar
2024

Enhanced Optical Character Recognition by Optical Sensor Combined with BERT and Cosine Similarity Scoring (Student Abstract)

AAAI 2024technical

Optical character recognition(OCR) is the technology to identify text characters embedded within images. Conventional OCR models exhibit performance degradation when performing with noisy images. To solve this problem, we propose a novel model, which combines computer vision using optical sensor wit…

Cited by 0SourcePDFScholar
2024

GRIL-Calib: Targetless Ground Robot IMU-LiDAR Extrinsic Calibration Method Using Ground Plane Motion Constraints

RA-L 2024

Targetless IMU-LiDAR extrinsic calibration methods are gaining significant attention as the importance of the IMU-LiDAR fusion system increases. Notably, existing calibration methods derive calibration parameters under the assumption that the methods require full motion in all axes. When IMU and LiD

Cited by 17SourcecodeScholar
2024

Microphone Pair Training for Robust Sound Source Localization With Diverse Array Configurations

RA-L 2024

We present a novel sound source localization method that leverages microphone pair training, designed to deliver robust performance in various real-world environments. Existing deep learning (DL)-based approaches face scalability issues when dealing with various types of microphone arrays. To addres

Cited by 10SourceScholar
2024

Virtual Action Actor-Critic Framework for Exploration (Student Abstract)

AAAI 2024technical

Efficient exploration for an agent is challenging in reinforcement learning (RL). In this paper, a novel actor-critic framework namely virtual action actor-critic (VAAC), is proposed to address the challenge of efficient exploration in RL. This work is inspired by humans' ability to imagine the pote…

Cited by 1SourcePDFScholar