← Search

Daechul Ahn

6 accepted papers

2026

BINDER: Instantly Adaptive Mobile Manipulation with Open-Vocabulary Commands

ICRA 2026poster

Open-vocabulary mobile manipulation (OVMM) requires robots to follow language instructions, navigate, and manipulate while updating their world representation as the environment changes dynamically. However, most prior works update their world representation only at discrete milestones, such as wayp…

2026

SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models

ICML 2026spotlight

Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic control, with test-time scaling (TTS) gaining attention to enhance robustness beyond training. However, existing TTS methods for VLAs require additional training, verifiers, and multiple forward pass…

Cited by 1SourceScholar
2025

ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO

AAAI 2025technical

Iterative self-improvement, a concept extending beyond personal growth, has found powerful applications in machine learning, particularly in transforming weak models into strong ones. While recent advances in natural language processing have shown its efficacy through iterative preference optimizati…

Cited by 0SourcePDFScholar
2024

Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

ACL 2024long

Recent advancements in large language models have influenced the development of video large multimodal models (VLMMs). Previous approaches for VLMMs involve Supervised Fine-Tuning (SFT) with instruction-tuned datasets, integrating LLM with visual encoders, and additional learnable parameters. Here,…

2023

Story Visualization by Online Text Augmentation with Context Memory

ICCV 2023poster

Story visualization (SV) is a challenging text-to-image generation task for the difficulty of not only rendering visual details from the text descriptions but also encoding a longterm context across multiple sentences. While prior efforts mostly focus on generating a semantically relevant image for…

Cited by 9PDFcodeScholar
2021

Zero-Shot Natural Language Video Localization

ICCV 2021poster

Understanding videos to localize moments with natural language often requires large expensive annotated video regions paired with language queries. To eliminate the annotation costs, we make a first attempt to train a natural language video localization model in zero-shot manner. Inspired by unsuper…

Cited by 62PDFcodeScholar