← Search

Gi-Cheon Kang

9 accepted papers

2025

CLIP-RT: Learning Language-Conditioned Robotic Policies from Natural Language Supervision

RSS 2025poster

Teaching robots desired skills in real-world environments remains challenging, especially for non-experts. Current robot learning methods often require expert demonstrations or complex programming, limiting their accessibility to non-experts. We posit that natural language offers an intuitive and ac…

Cited by 1PDFScholar
2025

Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following

ICRA 2025

Embodied Instruction Following (EIF) is the task of executing natural language instructions by navigating and interacting with objects in interactive environments. A key challenge in EIF is compositional task planning, typically addressed through supervised learning or few-shot in-context learning w

Cited by 8SourceScholar
2024

PGA: Personalizing Grasping Agents with Single Human-Robot Interaction

IROS 2024poster

Language-Conditioned Robotic Grasping (LCRG) aims to develop robots that comprehend and grasp objects based on natural language instructions. While the ability to understand personal objects like my wallet facilitates more natural interaction with human users, current LCRG systems only allow generic…

Cited by 2SourcecodeScholar
2024

PROGrasp: Pragmatic Human-Robot Communication for Object Grasping

ICRA 2024poster

Interactive Object Grasping (IOG) is the task of identifying and grasping the desired object via human-robot natural language interaction. Current IOG systems assume that a human user initially specifies the target object’s category (e.g., bottle). Inspired by pragmatics, where humans often convey t…

Cited by 7SourcecodeScholar
2023

GVCCI: Lifelong Learning of Visual Grounding for Language-Guided Robotic Manipulation

IROS 2023poster

Language-Guided Robotic Manipulation (LGRM) is a challenging task as it requires a robot to understand human instructions to manipulate everyday objects. Recent approaches in LGRM rely on pre-trained Visual Grounding (VG) models to detect objects without adapting to manipulation environments. This r…

Cited by 7SourcecodeScholar
2023

The Dialog Must Go On: Improving Visual Dialog via Generative Self-Training

CVPR 2023poster

Visual dialog (VisDial) is a task of answering a sequence of questions grounded in an image, using the dialog history as context. Prior work has trained the dialog agents solely on VisDial data via supervised learning or leveraged pre-training on related vision-and-language datasets. This paper pres…

2021

Attend What You Need: Motion-Appearance Synergistic Networks for Video Question Answering

ACL 2021long

Video Question Answering is a task which requires an AI agent to answer questions grounded in video. This task entails three key challenges: (1) understand the intention of various questions, (2) capturing various elements of the input video (e.g., object, action, causality), and (3) cross-modal gro…

2021

Reasoning Visual Dialog with Sparse Graph Learning and Knowledge Transfer

EMNLP 2021finding

Visual dialog is a task of answering a sequence of questions grounded in an image using the previous dialog history as context. In this paper, we study how to address two fundamental challenges for this task: (1) reasoning over underlying semantic structures among dialog rounds and (2) identifying s…

2020

Label Propagation Adaptive Resonance Theory for Semi-Supervised Continuous Learning

ICASSP 2020accepted

Semi-supervised learning and continuous learning are fundamental paradigms for human-level intelligence. To deal with real-world problems where labels are rarely given and the opportunity to access the same data is limited, it is necessary to apply these two paradigms in a joined fashion. In this pa…

Cited by 0SourceScholar