← Search

Byonghyo Shim

20 accepted papers

2026

Adaptive Capacity Allocation for Vision Language Action Fine-Tuning

ICRA 2026poster

Vision language action models (VLAs) are increasingly used for Physical AI, but deploying a pre-trained VLA model to unseen environments, embodiments, or tasks still requires adaptation. Parameter-efficient fine-tuning (PEFT), especially LoRA, is common for VLA policies, yet the exposed capacity kno…

2026

Mitigating Hallucination in Vision-Language Model with Depth and Spatial-aware Key-Value Refinement

ICLR 2026poster

Large vision–language models (VLMs) deliver state-of-the-art results on a wide range of multimodal tasks, yet they remain prone to visual hallucinations, producing content that is not grounded in the input image. Despite progress with visual supervision, reinforcement learning, and post-hoc attenti…

Cited by 0SourceScholar
2025

Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs

NeurIPS 2025spotlight

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view offers fine-grained cues about user attention and hand-object…

Cited by 0SourcecodeScholar
2025

Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models

ICLR 2025poster

Text-to-image generative models like DALL-E and Stable Diffusion have revolutionized visual content creation across various applications, including advertising, personalized media, and design prototyping. However, crafting effective textual prompts to guide these models remains challenging, often re…

2024

Expand-and-Quantize: Unsupervised Semantic Segmentation Using High-Dimensional Space and Product Quantization

AAAI 2024technical

Unsupervised semantic segmentation (USS) aims to discover and recognize meaningful categories without any labels. For a successful USS, two key abilities are required: 1) information compression and 2) clustering capability. Previous methods have relied on feature dimension reduction for informatio…

Cited by 1SourcePDFScholar
2024

Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models

EMNLP 2024finding

Recently, we have observed that Large Multi-modal Models (LMMs) are revolutionizing the way machines interact with the world, unlocking new possibilities across various multi-modal applications. To adapt LMMs for downstream tasks, parameter-efficient fine-tuning (PEFT) which only trains additional p…

2023

Depth-Relative Self Attention for Monocular Depth Estimation

IJCAI 2023poster

Monocular depth estimation is very challenging because clues to the exact depth are incomplete in a single RGB image. To overcome the limitation, deep neural networks rely on various visual hints such as size, shade, and texture extracted from RGB information. However, we observe that if such hints…

Cited by 5SourcePDFScholar
2023

Semantic-Aware Superpixel for Weakly Supervised Semantic Segmentation

AAAI 2023technical

Weakly-supervised semantic segmentation aims to train a semantic segmentation network using weak labels. Among weak labels, image-level label has been the most popular choice due to its simplicity. However, since image-level labels lack accurate object region information, additional modules such as…

2023

Semantic-Preserving Augmentation for Robust Image-Text Retrieval

ICASSP 2023accepted

Image-text retrieval is a task to search for the proper textual descriptions of the visual world and vice versa. One challenge of this task is the vulnerability to input image/text corruptions. Such corruptions are often unobserved during the training, and degrade the retrieval model’s decision qual…

Cited by 0SourceScholar
2023

Vision Transformer-Based Feature Extraction for Generalized Zero-Shot Learning

ICASSP 2023accepted

Generalized zero-shot learning (GZSL) is a technique to train a deep learning model to identify unseen classes using the image attribute. In this paper, we put forth a new GZSL technique exploiting Vision Transformer (ViT) to maximize the attribute-related information contained in the image feature.…

Cited by 0SourceScholar
2020

Deep Neural Network Based Matrix Completion for Internet of Things Network Localization

ICASSP 2020accepted

In this paper, we propose a deep neural network based matrix completion approach for Internet of Things (IoT) localization. In the proposed method, we recast Euclidean distance matrix completion problem into the alternating minimization problem. By using a cascade of multiple deep neural networks to…

Cited by 0SourceScholar
2018

A Compressive Sensing-Based Active User and Symbol Detection Technique for Massive Machine-Type Communications

ICASSP 2018accepted

In massive machine-type communication (mMTC) systems, a large number of machine-type devices sporadically transmit small packets with low rates. By exploiting the sporadic activity of machine-type devices, we can cast the detection problem as the compressive sensing-based multi-user detection (CS-MU…

Cited by 0SourceScholar
2016

Perfect error compensation via algorithmic error cancellation

ICASSP 2016accepted

This paper presents a novel statistical error compensation (SEC) technique — algorithmic error cancellation (AEC)-for designing robust and energy-efficient signal processing and machine learning kernels on scaled process technologies. AEC exhibits a perfect error compensation (PEC) property, i.e., i…

Cited by 0SourceScholar