← Search

YouHan Lee

4 accepted papers

2026

SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models

ICML 2026spotlight

Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic control, with test-time scaling (TTS) gaining attention to enhance robustness beyond training. However, existing TTS methods for VLAs require additional training, verifiers, and multiple forward pass…

Cited by 1SourceScholar
2024

BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining

AAAI 2024technical

The current research direction in generative models, such as the recently developed GPT4, aims to find relevant knowledge information for multimodal and multilingual inputs to provide answers. Under these research circumstances, the demand for multilingual evaluation of visual question answering (VQ…

Cited by 5SourcePDFScholar
2024

Pre-training Sequence, Structure, and Surface Features for Comprehensive Protein Representation Learning

ICLR 2024poster

Proteins can be represented in various ways, including their sequences, 3D structures, and surfaces. While recent studies have successfully employed sequence- or structure-based representations to address multiple tasks in protein science, there has been significant oversight in incorporating protei…

Cited by 17SourcePDFScholar
2022

Efficient Multilingual Multi-modal Pre-training through Triple Contrastive Loss

COLING 2022main

Learning visual and textual representations in the shared space from web-scale image-text pairs improves the performance of diverse vision-and-language tasks, as well as modality-specific tasks. Many attempts in this framework have been made to connect English-only texts and images, and only a few w…