← Search

Taewhan Kim

4 accepted papers

2025

CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation

CVPR 2025highlight

Correct use of electrical appliances has significantly improved human life quality. Unlike simple tools that can be manipulated with common sense, different parts of electrical appliances have specific functions defined by manufacturers. If we want the robot to heat bread by microwave, we should ena…

Cited by 0SourcePDFScholar
2025

ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?

IROS 2025

Visual actionable affordance has emerged as a transformative approach in robotics, focusing on perceiving interaction areas prior to manipulation. Traditional methods rely on pixel sampling to identify successful interaction samples or processing pointclouds for affordance mapping. However, these ap

Cited by 0SourcecodeScholar
2025

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning

AAAI 2025technical

Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP visual embedding. Because of this issue, there is a limitation that the image desc…

2024

IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning

EMNLP 2024main

Recent advancements in image captioning have explored text-only training methods to overcome the limitations of paired image-text data. However, existing text-only training methods often overlook the modality gap between using text data during training and employing images during inference. To addre…