← Search

Minyeong Kim

3 accepted papers

2026

D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI

ICLR 2026poster

Large language models leverage internet-scale text data, yet embodied AI remains constrained by the prohibitive costs of physical trajectory collection. Desktop environments---particularly gaming---offer a compelling alternative: they provide rich sensorimotor interactions at scale while maintaining…

Cited by 0SourcecodeScholar
2025

HalLoc: Token-level Localization of Hallucinations for Vision Language Models

CVPR 2025poster

Hallucinations pose a significant challenge to the reliability of large vision-language models, making their detection essential for ensuring accuracy in critical applications. Current detection methods often rely on computationally intensive models, leading to high latency and resource demands. The…

2024

Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models

ECCV 2024poster

"Hallucinations in vision-language models pose a significant challenge to their reliability, particularly in the generation of long captions. Current methods fall short of accurately identifying and mitigating these hallucinations. To address this issue, we introduce ESREAL, a novel unsupervised rei…

Cited by 5SourcePDFScholar