← Search

Minchan Kim

7 accepted papers

2026

D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI

ICLR 2026poster

Large language models leverage internet-scale text data, yet embodied AI remains constrained by the prohibitive costs of physical trajectory collection. Desktop environments---particularly gaming---offer a compelling alternative: they provide rich sensorimotor interactions at scale while maintaining…

Cited by 0SourcecodeScholar
2025

CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction

ICRA 2025

Real-life robot navigation involves more than just reaching a destination; it requires optimizing movements while addressing scenario-specific goals. An intuitive way for humans to express these goals is through abstract cues like verbal commands or rough sketches. Such human guidance may lack detai

Cited by 4SourceScholar
2025

Evidential-TTS: High Fidelity Zero-Shot Text-to-Speech Using Evidential Deep Learning

ICASSP 2025accepted

We propose Evidential-TTS, a novel zero-shot text-to-speech (TTS) system based on evidential deep learning (EDL). The model includes a length regulator to ensure precise alignment between phonemes and acoustic tokens. This module allows the evidential token generator to convert the aligned phoneme s…

Cited by 0SourceScholar
2024

Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models

ECCV 2024poster

"Hallucinations in vision-language models pose a significant challenge to their reliability, particularly in the generation of long captions. Current methods fall short of accurately identifying and mitigating these hallucinations. To address this issue, we introduce ESREAL, a novel unsupervised rei…

Cited by 5SourcePDFScholar
2023

EM-Network: Oracle Guided Self-distillation for Sequence Learning

ICML 2023poster

We introduce EM-Network, a novel self-distillation approach that effectively leverages target information for supervised sequence-to-sequence (seq2seq) learning. In contrast to conventional methods, it is trained with oracle guidance, which is derived from the target sequence. Since the oracle guida…

Cited by 3SourcePDFScholar
2023

Improving Learning Objectives for Speaker Verification from the Perspective of Score Comparison

ICASSP 2023accepted

Deep speaker embedding systems are usually trained with classification-based or end-to-end learning objectives. Popular end-to-end approaches utilize deep metric learning, which can be viewed as a few-shot classification objective. In this paper, we investigate the limit of conventional learning obj…

Cited by 0SourceScholar
2023

Pre-and Post-Contact Policy Decomposition for Non-Prehensile Manipulation with Zero-Shot Sim-To-Real Transfer

IROS 2023poster

We present a system for non-prehensile manipulation that require a significant number of contact mode transitions and the use of environmental contacts to successfully manipulate an object to a target location. Our method is based on deep reinforcement learning which, unlike state-of-the-art plannin…

Cited by 12SourceScholar